unikdahal commented on issue #6737:
URL: 
https://github.com/apache/datafusion-comet/issues/6737#issuecomment-6026242683

   Thanks @comphead, your diagram is right, and I worded it badly. iceberg-rust 
would still own the Iceberg semantics. I only meant: who owns the physical 
Parquet read.
   
   `IcebergTableScan` builds a `TableScan` and calls `to_arrow()`, with filters 
fixed into an Iceberg `Predicate` at plan time. Comet's `IcebergScanExec` also 
reads through `ArrowReader`. So neither gets DataFusion's ParquetSource runtime 
filter pruning (join/TopK), which is why #6641 currently needs equivalent 
support on the iceberg-rust ArrowReader path.
   
   By "pre-planned tasks" I meant the `FileScanTask`s Comet already gets from 
Iceberg Java. By "DataFusion-native source" I meant a `DataSource` behind 
`DataSourceExec`, so runtime filters have a standard place to land.
   
   So I agree with @mbutrovich: start with a `DataSource` that still reads 
through `ArrowReader`. Decoding through `ParquetSource` itself can be a later 
question. Deletes, especially equality deletes, are the hard part there.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to