+1 for moving to DataFusion PMC. The iceberg-datafusion integration is very important for Comet.
On Mon, Aug 31, 2026 at 11:42 PM Xuanwo <[email protected]> wrote: > > TBH, I also support moving to the DataFusion PMC. > > - DataFusion is the largest dependency in iceberg-datafusion. > - The largest downstream user of iceberg-datafusion is Comet, which shares > many of the same PMC members as DataFusion. > > It feels natural to be part of the DataFusion PMC. As long as DF PMC is > willing to accept this project, it LGTM. > > On Tue, Sep 1, 2026, at 12:03, Renjie Liu wrote: > > I don't think the testing should be a blocker of moving iceberg-datafusion > out of iceberg-rust repo. From what I learn, most of the sqllogictests are in > pr of modifying iceberg-datafusion integration, there are only few cases > where we rely on sqllogictests to verify features. > > > How would that sound for the `iceberg-rust` community to grant a couple of > > Apache DataFusion PMCs committer access to the repository, limited to the > > `iceberg-datafusion` crate via a CODEOWNERS file, so that they can help > > push reviews and PRs forward independently? > > I'm not sure if this is feasible, but moving the iceberg-datafusion crate to > apache datafusion project sounds a more reasonable approach to me. It's still > governed by Apache, and most of the code is related to DataFusion, so I think > the DataFusion community is in a better position to define the vision and > design for it. > > On Mon, Aug 31, 2026 at 10:45 PM Gabriel Musat <[email protected]> wrote: > > Hi everyone, > > Based on these two facts: > - There's a current reviewer bandwidth problem that hurts development > velocity of the `iceberg-datafusion` crate currently hosted under > `apache/iceberg-rust`. > - There is value in maintaining the `iceberg-datafusion` crate inside > `apache/iceberg-rust` for testing, design and governance. > > How would that sound for the `iceberg-rust` community to grant a couple of > Apache DataFusion PMCs committer access to the repository, limited to the > `iceberg-datafusion` crate via a CODEOWNERS file, so that they can help push > reviews and PRs forward independently? > > Based on review history, I'd propose Matt Butrovich and Tim Saucer as two > good candidates, but this would be completely up to the `iceberg-rust` > community. > > On 2026/08/26 17:56:19 Shawn Chang wrote: > > Hi all, > > > > Summarizing the discussion so far, including a few points raised in today’s > > community sync. > > > > There seems to be general agreement that the current DataFusion integration > > has a velocity/reviewer bandwidth problem, and moving it to a separate > > repository could help DataFusion contributors iterate more independently. > > At the same time, several open questions remain: > > > > - > > > > Whether repo separation is the right solution, versus expanding > > DataFusion reviewer/committer participation in iceberg-rust. > > - > > > > Where the boundary should be between engine specific integration and > > Iceberg core functionality, and how to avoid duplicated or forked Iceberg > > implementations. (how to avoid the case where Iceberg-datafusion moving > > much faster than the core and eventually need a forked core API > > implementation) > > - > > > > How compatibility and end-to-end correctness testing should work across > > repositories, since iceberg-rust still relies on DataFusion for > > integration > > testing. > > - > > > > Where the integration should live and how to keep it under Apache > > governance while making it easy for both Iceberg and DataFusion > > contributors to maintain. > > > > So I think the main question is not only whether to move the code, but how > > to improve development velocity without losing the close design, testing, > > and governance relationship between the engine integration and iceberg-rust > > core. > > > > > > Best, > > > > Shawn > > > > On Mon, Aug 24, 2026 at 4:14 AM Manu Zhang <[email protected]> wrote: > > > > > +1 moving the DataFusion integration and tests into a separate > > > apache-governed repository. Can we bring this discussion to the Community > > > Sync[1] this week? > > > > > > 1. > > > https://docs.google.com/document/d/1YuGhUdukLP5gGiqCbk0A5_Wifqe2CZWgOd3TbhY3UQg/edit?tab=t.0 > > > > > > On Mon, Aug 24, 2026 at 3:37 PM Renjie Liu <[email protected]> > > > wrote: > > > > > >> Hi, Matt: > > >> > > >> Thanks for raising this. > > >> > > >> 1) Apache governance: > > >> > > >> I would +1 for putting this in an apache repo, for example a sub repo of > > >> datafusion project. Shawn has stated most of the reasons, so I don't want > > >> to repeat it again. > > >> > > >> 2) Where do tests live/how tests should be maintained? > > >> > > >> Initially I was thinking about putting sqllogictest in iceberg-rust, but > > >> after second thought I'm leaning towards to put it in the new repo for > > >> two > > >> reasons: > > >> 1. It would be easier for developer of the datafusion-iceberg integration > > >> to add tests > > >> 2. It would make the dependency graph and version release easier. Though > > >> the dependency is on crate level rather than repo level, the bi-direction > > >> dependency may make version management weird and difficult. > > >> > > >> The downside of this approach is that it's a little unfriendly for > > >> iceberg-rust developers, but I think it's less frequent for iceberg-rust > > >> developers to add sqllogictests compared with datafusion-iceberg > > >> developers. > > >> > > >> 3) Dropping DataFusion dependencies from pyiceberg-core binding > > >> > > >> I'm not quite familiar with this part, but it sounds reasonable to me. > > >> > > >> > > >> On Sat, Aug 22, 2026 at 6:50 AM Shawn Chang <[email protected]> > > >> wrote: > > >> > > >>> Hi Matt, > > >>> > > >>> Thanks for raising this! I generally agree that we can move the > > >>> datafusion > > >>> integration to a separate repo if that helps more DataFusion experts to > > >>> work on the integrations > > >>> > > >>> On the three considerations: > > >>> 1) Apache governance: > > >>> I think this is my biggest concern so far. This could not only affect > > >>> contributors, but also whether the downstream users could continue to > > >>> use > > >>> the integration. > > >>> Even if the code remains Apache-licensed, governance, release of > > >>> artifacts, > > >>> and contribution policies can matter for adoption. > > >>> We should explore more about the option to keep it under an > > >>> Apache-governed > > >>> repository. > > >>> > > >>> 2) Where do tests live/how tests should be maintained? > > >>> > > >>> I think this is closely related to the governance question. To me, the > > >>> broader question is: *how do we make it easy for people from both the > > >>> Iceberg Rust and DataFusion communities to maintain the integration?* > > >>> > > >>> If we move it to a non-Apache repository, I worry that it could > > >>> eventually > > >>> look somewhat like the current Iceberg Java <> Trino integration: the > > >>> integration primarily lives on the Trino side and is therefore mostly > > >>> maintained by people who are already deeply involved in Trino. > > >>> > > >>> There is an important difference here, though. For Iceberg Java, Spark > > >>> is > > >>> arguably the primary engine integration and has a large Iceberg > > >>> contributor > > >>> base around it, so having the Trino integration maintained more > > >>> independently is relatively natural. In Iceberg Rust today, DataFusion > > >>> has > > >>> a much more central role. It is by far the most mature engine > > >>> integration > > >>> in the project, is used by our SQLLogicTest infrastructure, and many > > >>> users > > >>> building on Iceberg Rust are also building on DataFusion. The overlap > > >>> between the two communities is therefore much larger. > > >>> > > >>> Because of that, I would prefer that extracting the integration does not > > >>> turn it into something that is effectively owned only by the DataFusion > > >>> side. Ideally, both Iceberg Rust contributors and DataFusion > > >>> contributors > > >>> should be able to review changes, maintain compatibility, and evolve the > > >>> integration together. > > >>> > > >>> 3) Dropping DataFusion dependencies from pyiceberg-core binding: I think > > >>> it > > >>> makes sense to make things simpler and have left a comment on the github > > >>> issue with more detailed thoughts: > > >>> https://github.com/apache/iceberg-rust/issues/3036 > > >>> > > >>> Best, > > >>> Shawn > > >>> > > >>> On Fri, Aug 21, 2026 at 11:23 AM Matt Butrovich <[email protected]> > > >>> wrote: > > >>> > > >>> > Hello all, > > >>> > > > >>> > > > >>> > I've never tried emailing two different project lists at once, but it > > >>> was > > >>> > suggested that I do so to try to track the conversation across both > > >>> > communities. We'll see how this threads on the mailing lists. > > >>> > > > >>> > > > >>> > There has been a GitHub Discussion > > >>> > > > >>> https://github.com/apache/iceberg-rust/discussions/2992#discussioncomment-18082533 > > >>> , > > >>> > a GitHub Issue https://github.com/apache/iceberg-rust/issues/3029, and > > >>> > it's been a long topic of conversation in the past two weeks in both > > >>> the > > >>> > DataFusion and Iceberg Rust Community Calls to discuss moving the > > >>> > DataFusion integration from Iceberg Rust to a separate repository. I > > >>> will > > >>> > try to summarize some of the major points as I understand them, but > > >>> > the > > >>> > conversations are the ground truth and please feel free to correct me > > >>> here. > > >>> > > > >>> > > > >>> > The DataFusion TableProvider integration in Iceberg Rust makes > > >>> DataFusion > > >>> > a dependency for Iceberg Rust. The integration exists for multiple > > >>> reasons: > > >>> > > > >>> > 1) an engine to execute Iceberg Rust's corpus of sqllogictest files > > >>> > > > >>> > 2) a TableProvider integration for DataFusion users to interact with > > >>> > Iceberg tables > > >>> > > > >>> > > > >>> > Some of the motivations to break out the integration: > > >>> > > > >>> > 1) There have been a number of issues and pull requests against the > > >>> > DataFusion TableProvider in Iceberg Rust as users want to add more > > >>> > features, and they often go stale. I don't believe there are many > > >>> > committers/PMC members familiar with or using the DataFusion > > >>> integration. > > >>> > > > >>> > 2) Iceberg Rust would like to stay as engine-agnostic as possible. A > > >>> > recent DataFusion Ballista integration was declined for this reason > > >>> > https://github.com/apache/iceberg-rust/pull/2613. > > >>> > > > >>> > 3) Other projects that rely on both Iceberg Rust and DataFusion (e.g., > > >>> > Comet) are blocked by Iceberg Rust upgrading its DataFusion and Arrow > > >>> > dependencies before they can upgrade. > > >>> > > > >>> > > > >>> > Some considerations for both communities: > > >>> > > > >>> > 1) Where would this DataFusion TableProvider live? Most specifically, > > >>> > would it be an Apache-governed project? As folks like @andygrove point > > >>> out, > > >>> > this can affect whether some community members could contribute to it. > > >>> > There is a datafusion-contrib org for DataFusion-related projects to > > >>> have > > >>> > visibility but no Apache governance, but there may be options to put > > >>> > it > > >>> > under an Apache repository. > > >>> > > > >>> > 2) How would Iceberg Rust continue to run sqllogictests for regression > > >>> > testing? Does this live in a different repository that depends on this > > >>> new > > >>> > Iceberg Rust TableProvider crate? Would we be able to test Pull > > >>> Requests on > > >>> > Iceberg Rust with sqllogictests? > > >>> > > > >>> > 3) @kevinqliu is familiar with the Python bindings in Iceberg Rust, > > >>> and I > > >>> > believe that DataFusion dependency might be removed as well, but that > > >>> is > > >>> > not the core focus of this conversation. He has also proposed removing > > >>> that > > >>> > and using the DataFusion Python bindings directly. > > >>> > > > >>> > > > >>> > I'm sure I'm forgetting things, but this email is long enough. > > >>> > > > >>> > > > >>> > Thanks everyone for the discussion thus far, and looking forward to > > >>> more > > >>> > input on this. > > >>> > > > >>> > > > >>> > -Matt > > >>> > > > >>> > > > >>> > > > >>> > > > >>> > > >> > > > > Xuanwo > > https://xuanwo.io/ >
