mohitgurav20 commented on issue #6418:
URL: 
https://github.com/apache/datafusion-comet/issues/6418#issuecomment-5896751865

   Looking into `q68` and `q54`, a couple of things stand out when comparing 
Spark 3.5 vs 4.2 behavior:
   
   ### `q68` slowdown (1.42s -> 3.94s on Spark 4.2)
   `q68` and `q46` share almost identical 5-table star join structures, but 
`q68`'s date filter (`d_dom BETWEEN 1 AND 2`) is significantly more selective 
(~6.6% of `store_sales`) than `q46`'s `d_dow IN (6,0)` (~28.5%).
   
   Because of that high selectivity at SF1000, Spark 4.2 is very likely 
triggering runtime Bloom filter injection (`InjectRuntimeFilter`). Since 
runtime Bloom filters currently fall back on Spark 4.2 (#4968), Comet ends up 
falling back to Spark right while reading `store_sales`. That inserts heavy 
`ColumnarToRow` / `RowToColumnar` conversions across a massive table scan, 
which would explain why `q68` regressed while `q46` (which doesn't trigger the 
filter) actually got faster.
   
   **Quick test to confirm**: run `q68` with 
`spark.sql.optimizer.runtimeFilter.enabled=false` on Spark 4.2 and see if 
execution time drops back to ~1.5s.
   
   ### `q54` baseline gap (2.54s Spark vs 5.91s Comet)
   `q54` has 3 nested aggregation CTEs with scalar subqueries on `date_dim`. On 
Spark 4.2, `MergeSubplans` (#5834) merges those scalar subqueries into a single 
struct-returning subquery. Because Comet doesn't support struct-typed scalar 
subqueries yet, the projection that consumes it falls back to JVM Spark. In a 
multi-stage query like `q54`, falling back mid-pipeline breaks native execution 
and bounces data back and forth between Rust and the JVM.
   
   **Quick test to confirm**: run `q54` with 
`spark.sql.optimizer.mergeScalarSubqueries.enabled=false`.
   
   If the event logs from the SF1000 run are available, checking for 
`ColumnarToRow` transitions on `store_sales` in `q68` or around the merged 
subquery projection in `q54` should confirm this right away.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to