rich7420 opened a new pull request, #6396: URL: https://github.com/apache/datafusion-comet/pull/6396
## Which issue does this PR close? Part of #6093. Extracts the generic filter output optimization from #6180, following [review feedback](https://github.com/apache/datafusion-comet/pull/6180#discussion_r4127884161). This PR does not close the native AtLeastNNonNulls issue. ## Rationale for this change A native filter currently materializes every input column even when its parent only consumes a subset, or only needs the row count. Wide inputs with narrow projections spend time filtering arrays whose values are never read. ## What changes are included in this PR? Push distinct required column indices from a column-only projection into DataFusion's filter output projection. Keep the parent projection to restore aliases, order and duplicate outputs. The predicate retains its original input schema, and both native plans remain under the existing Spark metrics nodes, including shared plan IDs. Empty projections preserve row counts. Computed projections and projections requiring every input column retain their existing behavior. ## How are these changes tested? - [Fork CI at `256239ebe`](https://github.com/rich7420/datafusion-comet/actions/runs/36542579994) passed native build, Rust tests, Spark 4.1 Comet suites, TPC-H/TPC-DS result checks and lint. Spark's own 4.1 SQL suites also passed: Catalyst, all three SQL Core shards and all three Hive shards. - Locally, 50 Rust planner tests and both focused Spark SQL/metrics tests passed. Planner assertions check the actual filter projection, output mapping, zero-column row counts and shared/distinct plan IDs. SQL cases cover reordered/duplicate outputs, aliases, NULLs, empty results and a rejected row containing an invalid ANSI cast. The Scala test checks native filter/project metrics. - Prior measurements of this pruning implementation reduced narrow-output string query time by about 19–22%, with full-output controls unchanged. Those measurements used the #6180 workload; they are not fresh measurements of this extracted branch. The branch is based on `ba9aa33db`. A merge-tree check against upstream main `8369bf11a` is conflict-free and retains the same implementation and test files. New main changes overlap only in a different section of the operator guide. Upstream CI will validate that merge result. Other Spark-profile SQL suites and Iceberg suites were not run for this candidate. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
