I agree with Russell here. There's already a proposal (though it needs to be updated), so I think we should consolidate around updating that as opposed to starting a new effort around this.
Original thread is here: https://lists.apache.org/thread/kz09b2rj7c00j6z2vlqg8v5myh94bgl5 -Dan On Tue, Jul 28, 2026 at 11:41 AM Russell Spitzer <[email protected]> wrote: > Hopefully bumping this out of Spam. I know Talat and company have a > proposal as well and it would be great to see some consensus amongst these > proposals but it seems like everyone is on vacation at the moment. > > > On Sun, Jul 26, 2026 at 10:21 AM Nitya Kumar Sharma < > [email protected]> wrote: > >> Hi all, >> >> >> Now that the Parquet FILE logical type has merged (parquet-format#585 >> [1]), I'd like to share a proposal for a corresponding file data type in >> Iceberg — >> https://docs.google.com/document/d/1W25ztRqm2xXBxOzGtUHK_mv-FLviQeIcKnUcmvCKOuo/edit >> — and start a discussion focused on the type-level semantics. >> >> >> A note on scope and framing: this document deliberately covers only the >> data type (the value) — how file lives in the type system, schema >> serialization/evolution, statistics, and behavior under >> transforms/deletes/time-travel. It leaves credential vending and brokered >> access to the catalog. >> >> >> I want to be explicit that this is not a counter-proposal to the existing >> document from Talat Uyarer, Alex Stephen, and Sung Yun [2], which covers >> both the data type and catalog-brokered credential vending. I've scoped >> this one narrowly on purpose, so we can converge on the type semantics >> independently of the access-control layer — which can sit on top unchanged, >> including their REST/credential-vending design. My intent is to keep the >> discussion moving and focused on the file data type, and I'm happy to >> collaborate and fold this into a shared direction. >> >> >> A few Iceberg-specific decisions in the doc, for discussion: >> >> >> >> - >> >> Representation: file is a "named struct" — it serializes with a "file" >> discriminator and a fixed, closed set of six leaves (uri, offset, size, >> content_type, checksum, inline), each assigned a normal field ID. >> Field semantics, byte resolution, and checksum format defer to #585. >> - >> >> Two forms only (inline and external); no in-file/self-reference: #585 >> permits packing bytes within the same data file and addressing them by >> offset/size with no uri. This proposal does not adopt that in-file >> form — only inline bytes and external uri references are allowed — >> keeping the type to two well-understood forms while leaving room to add it >> later. >> - >> >> Relative URIs resolve against the table's base location (v4 >> relative-location rules), so tables stay relocatable; #585 permits >> relative >> URIs but leaves "relative to what" undefined. >> - >> >> Statistics use ordinary field-ID-keyed bounds per leaf (bounds on uri >> enable object-prefix file skipping) — no new bounds structure, in contrast >> to variant. >> - >> >> Not comparable as a whole: void transform only, and file can't be a >> partition/sort or equality-delete key. >> - >> >> Credentials/access are out of scope for the type (a catalog concern). >> >> >> Feedback very welcome — especially on the representation (the "named >> struct" question). >> >> >> Thanks, Nitya >> >> >> [1] https://github.com/apache/parquet-format/pull/585 >> >> [2] >> https://docs.google.com/document/d/1BdA4G-38GoVcpX6UY4UI-QA1MvrDnH6uSKeCJBadILQ/edit >> >> >>
