andygrove commented on PR #5533:
URL: 
https://github.com/apache/datafusion-comet/pull/5533#issuecomment-5876573343

   This is a light fully automated review since there are so many PRs open.
   
   When a filter that decodes sits under a limit, this also takes the Parquet 
scan out of native execution. Spark's `FileSourceStrategy` copies every 
deterministic data-column predicate into `dataFilters`, including ones it can't 
push into Parquet, so for the fixture query at `unbase64_operator_masks.sql:35` 
the scan's own `dataFilters` hold `unbase64(bad) <=> X'616263'`. `originalPlan` 
at `CometExecRule.scala:892` exposes them, `limitName` at line 968 matches on 
the scan, and line 1010 rebuilds it as a Spark `FileSourceScanExec`. As far as 
I can tell the native reader never decodes there by default. Per the comment at 
`native/core/src/parquet/parquet_exec.rs:201`, data filters are only evaluated 
per row when `spark.comet.parquet.rowFilterPushdown.enabled` is on, and that 
defaults to `false`. Otherwise they only feed row-group, page-index and 
bloom-filter pruning, and DataFusion's pruning rewrite rejects a function call 
over a column. The Spark `Filter` above, which already stays in 
 Spark, is what decodes row by row. Could the scan only count as an evaluation 
site when row filter pushdown is enabled? `checkSparkAnswerAndFallbackReason` 
passes either way, so a `nativeScans(plan) == 1` check on the same query at 
`CometEvaluationMaskSuite.scala:192`, like the one at line 235, would pin it.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to