andygrove opened a new pull request, #6309: URL: https://github.com/apache/datafusion-comet/pull/6309
## Which issue does this PR close? No issue. This is a documentation refresh of the contributor roadmap, following the last pass in #5064. ## Rationale for this change Several roadmap sections still describe work as unstarted or blocked when it has since landed, and a few references point at issues that have closed or code that has moved. Contributors use the roadmap to find where work is coordinated, so stale entries send them to the wrong place. ## What changes are included in this PR? Only `docs/source/contributor-guide/roadmap.md` changes. Each section was checked against main and the linked issues and PRs: - **Iceberg Table Writes**: the section said no design was committed. The split writer/committer plan (#4658) and the iceberg-rust writer (#5361) have landed behind two experimental, off-by-default settings. The section now describes them, notes that merge-on-read writes aren't intercepted yet (#6240), and points at the production-quality epic (#5649) and the default flip (#5644). The closed draft #4487 is dropped. - **Iceberg Table Format V3 Support**: native deletion vector reads landed in #5853, which replaced the closed draft #4887, and the upstream iceberg-rust epic has closed. The section now lists what still falls back: row lineage metadata columns, column default values, and the `variant`/`geometry`/`geography`/`unknown` types. The HDFS note referred to a scheme match in `iceberg_scan.rs` that has since moved to `iceberg_common.rs`. It now describes the supported storage backends without naming a file. - **Native Coverage for Codegen-Dispatched Expressions**: the section said aggregate functions go through the codegen-dispatch bridge. The dispatcher only handles scalar expressions (`CometBatchKernelCodegen` rejects `AggregateFunction`), so incompatible aggregates fall back to Spark. The link now points at the expression reference, whose Implementation column lists the dispatched expressions; the compatibility guide it pointed at doesn't have that list. - **Java/Scala UDF Support**: the section said Python UDFs always fall back. `mapInArrow` and `mapInPandas` have had experimental support since #4234. - **TPC-H and TPC-DS Performance**: the TPC-DS epic #858 is closed, so the section now points at #2551. - **Upstream Work in DataFusion**: nearly every `datafusion-spark` function now has a Comet serde (#4150); the remaining exception, `json_tuple`, is tracked in #3160. - **Spillable Hash Join**: adds links to #2545 and to the upstream DataFusion design for spilling `HashJoinExec` (apache/datafusion#24768). - **Delta Lake Support**: describes the delta-spark contrib scan in review (#5365) and the proposal to converge it with the `delta-kernel-rs` path (#5411). The dormant plain-table draft #4669 is dropped because #5365 supersedes its approach; I can add it back if we'd rather keep it listed. The window, lambda, native Parquet write, and memory management sections are still accurate and are unchanged. ## How are these changes tested? This is a documentation-only change. `prettier --check` passes on the file, and every reference-style link it uses is defined, with no unused definitions. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
