unikdahal opened a new issue, #6737:
URL: https://github.com/apache/datafusion-comet/issues/6737
I've been working on runtime pruning for `IcebergScanExec` (#6588, #6641,
follow-up to #4807/#4808). It works. On the benchmark in #6641, the sorted join
and TopK cases cut reader bytes by 94-99%. But getting there meant adding a
fair amount of adaptive-pruning machinery to iceberg-rust (tracked in
apache/iceberg-rust#3343). That includes a runtime predicate provider,
rejecting file tasks from retained stats, live row-group refresh, and keeping
position/equality deletes correct through all of it.
While writing it I kept noticing that DataFusion's Parquet path already does
most of this for plain Parquet scans. apache/datafusion#22450 (closing #22407)
re-checks a live `DynamicFilterPhysicalExpr` at row-group boundaries, and the
sort pushdown epic (apache/datafusion#23036) covers stats-based file/RG
ordering for TopK. So for the same dynamic filters we end up with two parallel
stacks:
```
native Parquet scan: DynamicFilterPhysicalExpr -> ParquetSource -> file /
RG / page pruning
Iceberg scan: DynamicFilterPhysicalExpr -> translate -> iceberg-rust
runtime predicate
-> ArrowReader -> file / RG / page pruning
(reimplemented)
```
To be clear, I'm not proposing to rip anything out, and #6641 stands as is.
I also don't want Comet re-planning tables in Rust. Iceberg Java should keep
producing the `FileScanTask`s.
What I'm wondering about is `apache/datafusion-iceberg`, now that it lives
under DataFusion. Long term, could it expose a source that takes
already-planned `FileScanTask`s and runs them through
`DataSourceExec`/`ParquetSource`, with the Iceberg-specific parts wrapped
around it (field-ID schema mapping, defaults, position/equality deletes, DVs,
encryption)?
```
Spark / Iceberg Java -> FileScanTask[] -> Comet -> datafusion-iceberg
-> DataSourceExec / ParquetSource -> Parquet
```
Comet would hand it the Java-planned tasks, and dynamic filters, TopK
ordering and whatever DataFusion adds next would come for free instead of being
redone for Iceberg.
I know the hard part is semantics. Position deletes need physical row
positions before any filtering, equality deletes need extra columns pulled into
the internal projection, and then there are DVs, split tasks and schema
evolution. `ArrowReader` handles all of that today and Comet relies on it.
There's also a smaller practical point in #6126: for the native Parquet scan
Comet owns the `AsyncFileReader`, but for Iceberg it's iceberg-rust's
`ArrowFileReader` and we can't swap it, which makes things like memory
accounting harder.
## Questions
1. Is this a direction people would want, or is `ArrowReader` meant to stay
the physical execution path for Iceberg in Comet?
2. If it is, is datafusion-iceberg the right home, and would a "from
pre-planned tasks" entry point be reasonable there?
3. Meanwhile, is it fine to keep going with #6641 in its current shape,
knowing parts of it might be replaced later? I mainly want to decide how much
more adaptive logic (live refresh, ordering) to put into iceberg-rust.
I haven't dug into datafusion-iceberg's internals yet, so some of this may
already exist or be planned.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]