That sounds a lot more sustainable long term, so I'm excited to see the proposal.
On Tue, Aug 18, 2026 at 10:22 PM Alkis Evlogimenos < [email protected]> wrote: > I am withdrawing this. PR 603 [1] now removes self-references from FILE > instead of specifying compression and encryption for them. > > Through extensive discussion with folks we think that a non-contiguous > page mechanism belongs to the non-contiguous page proposal and not as a > special way to do FILE. This means we should disallow self-references in > FILE and let the non-contiguous pages proposal to allow storing outside > values of the `inline` column elsewhere in the file. If those values are in > pages they naturally get al the fields: `is_compressed`, `encryption`, > `uncompressed_page_size` etc. As such non-contiguous pages composes with > FILE instead of competing with it. > > We will send the non-contiguous pages proposal in a separate thread. > > Cheers, > > [1] https://github.com/apache/parquet-format/pull/603 > > > On Thu, Jul 30, 2026 at 9:55 AM Alkis Evlogimenos < > [email protected]> wrote: > >> Great! >> >> I have updated the PR [1] with precise semantics and added encryption >> rules. >> >> The remaining question is the uncompressed size. I did a deeper analysis >> of the codecs. Snappy includes it in its compressed representation, but >> several other Parquet codecs either omit it or do not guarantee it. The >> current proposal requires readers to support dynamically sized >> decompression output. I think that's an acceptable tradeoff but I am open >> to alternatives if there is strong need to know the uncompressed size >> before decompression. >> >> [1] https://github.com/apache/parquet-format/pull/603 >> >> On Thu, Jul 30, 2026 at 3:45 AM Russell Spitzer < >> [email protected]> wrote: >> >>> I think that's pretty accurate. We should probably also establish >>> encryption rules concurrently with any compression rules we add. >>> >>> If I can summarize the pro argument: >>> >>> "A Parquet writer is already writing these bytes; had they been inlined >>> they would have had compression applied. Therefore, if a writer is writing >>> a blob into non-columnar space, it should use the same compression as it >>> would have if it were writing into an inline page." >>> >>> My initial hesitation was that I thought the consumer of this object >>> would want to interact with the blob using only the offset. Daniel's >>> perspective is that the Parquet reader will always be an interlocutor for >>> getting these bytes, so we can rely on inheritance from Parquet internals >>> to produce the actual byte stream. That clarifies the model, and it's also >>> why I'm more in favor of an explicit definition here. Inheritance means we >>> cannot use the offset on its own, so we essentially have three modes of >>> reading: >>> >>> Inline - The Parquet reader returns the bytes in the type object. >>> Self-reference - The reader could return offset/size, but that may be >>> meaningless or miss compression information. If an engine wants the logical >>> bytes, it needs a byte stream from the Parquet reader and can't just copy >>> the range as-is from the file. >>> External reference - The reader returns a path; it's up to the consumer >>> to know what to do with it. >>> >>> I thought we were basically doing self reference and external reference >>> the same way: whether there is compression or not is a function of the >>> actual file contents, not of the fact that the range lives inside a Parquet >>> file. I can see the benefit of inheriting, but in that case I think we >>> should treat self-ref bytes like other writer-owned storage (compression, >>> encryption, and anything similar we add later), and not just add only on >>> codec inheritance. I think the pro inheritance case is basically saying we >>> should have inline and self-reference behave the same with external >>> reference as the outlier. >>> >>> On the open points in (1): I'm still uneasy about living with the >>> snappy/lz4_raw uncompressed-size gap, and about tying is_compressed to "the >>> page containing the self-reference" when the blob itself is out-of-band. >>> Those feel like more reasons to prefer a small explicit layout over a >>> growing set of inheritance rules. But if that's what the majority wants to >>> go with I don't have a problem with it. >>> >>> >>> On Wed, Jul 29, 2026 at 6:19 PM Alkis Evlogimenos via dev < >>> [email protected]> wrote: >>> >>>> This was discussed at the Parquet Sync tonight. >>>> >>>> There is general agreement that compression is valuable for >>>> self-referenced >>>> data like large text blobs and that external references should remain >>>> outside of Parquet's compression model. There is also agreement that we >>>> need to address this before releasing the format as it renders >>>> self-references impractical for text data. We should resolve this before >>>> releasing the change to the format, while it can be changed cheaply. >>>> >>>> >>>> The two positions are: >>>> >>>> 1. inherited compression >>>> >>>> The Parquet writer owns the bytes written to the file and applies the >>>> same >>>> compression decision whether a value is stored inline or as a >>>> self-reference. >>>> >>>> This avoids adding per-value compression metadata, but the specification >>>> must still define: >>>> a. how reader obtains uncompressed size: all compressors except snappy >>>> and >>>> lz4raw provide this natively -> I suggest we can live with that gap. >>>> b. how the compression decision is associated with a self-reference when >>>> is_compressed is set to false in Data Page V2 -> I suggest we inherit >>>> the >>>> is_compressed from the page containing the self-reference. This provides >>>> flexibility to the writer to adjust the compression decision mid-stream. >>>> c. what happens when the FILE schema does not contain an inline field >>>> -> I >>>> suggest that it defaults to RAW. >>>> >>>> 2. explicit compression metadata/framing >>>> >>>> Each self-reference explicitly records its compression information, >>>> either >>>> through additional fields or through a framing format. This is >>>> self-describing and allows compression to vary between values, but adds >>>> metadata and format complexity compared with the inheritance rule >>>> proposed >>>> in [1]. >>>> >>>> >>>> Russel does this capture the discussion accurately? >>>> >>>> My preference remains inheritance. It makes the ownership model >>>> stronger: >>>> the Parquet writer owns both the inline and self-reference >>>> representations >>>> and avoids exposing a second compression policy to users. I agree that >>>> PR >>>> needs some more refinement, I will update it shortly. >>>> >>>> >>>> >>>> > Let's not rush last-minute additions after the vote ended? >>>> >>>> Compression inheritance was part of the proposal that was voted on and >>>> was >>>> removed at the last minute [2]. Regardless of how we think about that, I >>>> agree we should not rush an underspecified change. At the same time we >>>> should focus on resolving compressibility of text data before the first >>>> release of FILE because it will be substantially easier to do so >>>> compared >>>> to changing it afterward. >>>> >>>> [1] https://github.com/apache/parquet-format/pull/603 >>>> [2] https://lists.apache.org/thread/qmx5vxg8y76xxx90cqlcfvrj7d25ps6s >>>> >>>> Cheers, >>>> >>>> On Wed, Jul 29, 2026 at 10:42 PM Antoine Pitrou <[email protected]> >>>> wrote: >>>> >>>> > >>>> > I agree with Russell here. Let's not rush last-minute additions after >>>> > the vote ended? >>>> > >>>> > Regards >>>> > >>>> > Antoine. >>>> > >>>> > >>>> > Le 29/07/2026 à 19:17, Russell Spitzer a écrit : >>>> > >> >>>> > >> This is equivalent to any data written in parquet and compressed >>>> as a >>>> > >> page, so I don't see the issue there. >>>> > > >>>> > > >>>> > > This is the problem I have. We haven't defined how the Parquet >>>> writes >>>> > these >>>> > > bytes, so assuming the self reference can be treated as a page seems >>>> > > undefined to me. >>>> > > >>>> > > On Wed, Jul 29, 2026 at 12:12 PM Alkis Evlogimenos via dev < >>>> > > [email protected]> wrote: >>>> > > >>>> > >> I opened a PR for this change here: >>>> > >> https://github.com/apache/parquet-format/pull/603 >>>> > >> >>>> > >> On Wed, Jul 29, 2026 at 7:49 PM Daniel Weeks <[email protected]> >>>> wrote: >>>> > >> >>>> > >>> Russell, I'm not sure I follow your points here. With inline, >>>> the data >>>> > >> is >>>> > >>> compressed using the column compression defined in the writer. >>>> This is >>>> > >>> equivalent to any data written in parquet and compressed as a >>>> page, so >>>> > I >>>> > >>> don't see the issue there. >>>> > >>> >>>> > >>> The self-ref/in-file offset+size would represent the compressed >>>> size >>>> > >> (when >>>> > >>> compressed). Having the uncompressed size is nice, but >>>> technically not >>>> > >>> guaranteed to be accurate because uncompressed sizes rely heavily >>>> on >>>> > the >>>> > >>> memory layout which can vary by language/implementation. >>>> > >>> >>>> > >>> I think this approach largely aligns with how parquet handles data >>>> > >> managed >>>> > >>> by the writer. >>>> > >>> >>>> > >>> -Dan >>>> > >>> >>>> > >>> On Wed, Jul 29, 2026 at 9:37 AM Russell Spitzer < >>>> > >> [email protected] >>>> > >>>> >>>> > >>> wrote: >>>> > >>> >>>> > >>>> I'm a little worried about adding this in, >>>> > >>>> >>>> > >>>> Self-references are already specified as [offset, offset+size) >>>> ranges: >>>> > >>> the >>>> > >>>> resolved bytes are that range, and no compression transform is >>>> > defined. >>>> > >>>> Adding compression isn't a one-line codec inheritance rule. If we >>>> > reuse >>>> > >>>> Parquet's CompressionCodec model, readers also need framing (at >>>> least >>>> > >>>> uncompressed length). "The inline column chunk's >>>> CompressionCodec" is >>>> > >>>> underspecified too: a FILE group may omit inline, codecs are per >>>> > column >>>> > >>>> chunk / row group, and v2 can leave data uncompressed under the >>>> same >>>> > >>> chunk >>>> > >>>> codec via is_compressed. >>>> > >>>> >>>> > >>>> Encryption is a good example of the same gap. The merged text >>>> says >>>> > >>> self-ref >>>> > >>>> files must not use modular encryption, but we never really worked >>>> > >> through >>>> > >>>> how those ranges would interact with features that need >>>> page/module >>>> > >>>> structure. We shouldn't bolt compression onto the same ranges >>>> without >>>> > >>>> defining how compressed self-ref bytes are laid out. >>>> > >>>> >>>> > >>>> IMHO, we should first strictly define how these out-of-band >>>> bytes are >>>> > >>> laid >>>> > >>>> out and framed, then we can make more concrete decisions about >>>> > >> inheriting >>>> > >>>> codecs, encryption, and so on. >>>> > >>>> >>>> > >>>> On Wed, Jul 29, 2026 at 9:11 AM Alkis Evlogimenos via dev < >>>> > >>>> [email protected]> wrote: >>>> > >>>> >>>> > >>>>> Hello, >>>> > >>>>> >>>> > >>>>> Now that FILE [1] is merged I'd like to reopen one point from >>>> the >>>> > >>>> original >>>> > >>>>> proposal that got dropped before merge: the bytes of a >>>> self-reference >>>> > >>>>> should use the same CompressionCodec as the column's inline >>>> field. >>>> > >>>> Removing >>>> > >>>>> it was a mistake, and it's cheap to fix while FILE hasn't >>>> shipped in >>>> > >> a >>>> > >>>>> release. >>>> > >>>>> >>>> > >>>>> As merged, a self-reference can only be stored uncompressed. >>>> That's >>>> > >> ok >>>> > >>>> for >>>> > >>>>> data like images and video, but it makes self-references >>>> useless for >>>> > >>> text >>>> > >>>>> blobs (html, json, logs). Values too big to be inline yet small >>>> > >> enough >>>> > >>> to >>>> > >>>>> want as self-references are exactly where PLAIN blows up the >>>> storage >>>> > >>>> cost. >>>> > >>>>> A self-reference is the Parquet writer's decision to store a >>>> large >>>> > >>> value >>>> > >>>>> out-of-band, so it should be compressed consistently with the >>>> inline >>>> > >>>> values >>>> > >>>>> it came from. >>>> > >>>>> >>>> > >>>>> To the objections from the PR: >>>> > >>>>> >>>> > >>>>> 1. Apply it uniformly to external refs too. >>>> > >>>>> >>>> > >>>>> The asymmetry is the point. An external s3://… can be >>>> referenced by >>>> > >>> many >>>> > >>>>> files and systems, including ones that know nothing of Parquet; >>>> its >>>> > >>>>> encoding is decided above Parquet, sometimes outside any >>>> engine. A >>>> > >>>>> self-reference lives inside Parquet and is written by the >>>> Parquet >>>> > >>> writer. >>>> > >>>>> Parquet owns those bytes, so Parquet compresses them. >>>> > >>>>> >>>> > >>>>> 2. Let the engine own it. >>>> > >>>>> >>>> > >>>>> The engine already owns the blob's own compression via >>>> content_type >>>> > >> and >>>> > >>>> can >>>> > >>>>> pass the writer pre-compressed bytes. The inline codec is a >>>> different >>>> > >>>>> thing: the storage compression Parquet applies to the column. A >>>> > >>>>> self-reference is the same bytes spilled out-of-band, so it >>>> belongs >>>> > >> to >>>> > >>>> the >>>> > >>>>> same storage and codec. >>>> > >>>>> >>>> > >>>>> 3. Force-compressing images wastes CPU. >>>> > >>>>> >>>> > >>>>> It doesn't. The writer picks the codec per column chunk (and >>>> per page >>>> > >>> in >>>> > >>>>> v2), same as it already does for inline. Inheritance just >>>> propagates >>>> > >>> that >>>> > >>>>> choice. A column chunk written uncompressed stays uncompressed. >>>> > >>>>> >>>> > >>>>> 4. Compaction would force decompress/recompress. >>>> > >>>>> >>>> > >>>>> Only an issue for external refs. Compaction rewrites the whole >>>> file, >>>> > >> so >>>> > >>>>> self-referenced bytes re-encode in the same pass, exactly like >>>> inline >>>> > >>>>> values. >>>> > >>>>> >>>> > >>>>> Proposed wording: >>>> > >>>>> >>>> > >>>>>> The bytes referenced by a self-reference (a FILE with no uri) >>>> are >>>> > >>>>> compressed with the CompressionCodec of the inline column >>>> chunk's >>>> > >>>>> ColumnMetadata. This does not apply to external references. >>>> > >>>>> >>>> > >>>>> Cheers, >>>> > >>>>> >>>> > >>>>> [1] https://github.com/apache/parquet-format/pull/585 >>>> > >>>>> >>>> > >>>> >>>> > >>> >>>> > >> >>>> > > >>>> > >>>> > >>>> > >>>> >>>
