I am withdrawing this. PR 603 [1] now removes self-references from FILE instead of specifying compression and encryption for them.
Through extensive discussion with folks we think that a non-contiguous page mechanism belongs to the non-contiguous page proposal and not as a special way to do FILE. This means we should disallow self-references in FILE and let the non-contiguous pages proposal to allow storing outside values of the `inline` column elsewhere in the file. If those values are in pages they naturally get al the fields: `is_compressed`, `encryption`, `uncompressed_page_size` etc. As such non-contiguous pages composes with FILE instead of competing with it. We will send the non-contiguous pages proposal in a separate thread. Cheers, [1] https://github.com/apache/parquet-format/pull/603 On Thu, Jul 30, 2026 at 9:55 AM Alkis Evlogimenos < [email protected]> wrote: > Great! > > I have updated the PR [1] with precise semantics and added encryption > rules. > > The remaining question is the uncompressed size. I did a deeper analysis > of the codecs. Snappy includes it in its compressed representation, but > several other Parquet codecs either omit it or do not guarantee it. The > current proposal requires readers to support dynamically sized > decompression output. I think that's an acceptable tradeoff but I am open > to alternatives if there is strong need to know the uncompressed size > before decompression. > > [1] https://github.com/apache/parquet-format/pull/603 > > On Thu, Jul 30, 2026 at 3:45 AM Russell Spitzer <[email protected]> > wrote: > >> I think that's pretty accurate. We should probably also establish >> encryption rules concurrently with any compression rules we add. >> >> If I can summarize the pro argument: >> >> "A Parquet writer is already writing these bytes; had they been inlined >> they would have had compression applied. Therefore, if a writer is writing >> a blob into non-columnar space, it should use the same compression as it >> would have if it were writing into an inline page." >> >> My initial hesitation was that I thought the consumer of this object >> would want to interact with the blob using only the offset. Daniel's >> perspective is that the Parquet reader will always be an interlocutor for >> getting these bytes, so we can rely on inheritance from Parquet internals >> to produce the actual byte stream. That clarifies the model, and it's also >> why I'm more in favor of an explicit definition here. Inheritance means we >> cannot use the offset on its own, so we essentially have three modes of >> reading: >> >> Inline - The Parquet reader returns the bytes in the type object. >> Self-reference - The reader could return offset/size, but that may be >> meaningless or miss compression information. If an engine wants the logical >> bytes, it needs a byte stream from the Parquet reader and can't just copy >> the range as-is from the file. >> External reference - The reader returns a path; it's up to the consumer >> to know what to do with it. >> >> I thought we were basically doing self reference and external reference >> the same way: whether there is compression or not is a function of the >> actual file contents, not of the fact that the range lives inside a Parquet >> file. I can see the benefit of inheriting, but in that case I think we >> should treat self-ref bytes like other writer-owned storage (compression, >> encryption, and anything similar we add later), and not just add only on >> codec inheritance. I think the pro inheritance case is basically saying we >> should have inline and self-reference behave the same with external >> reference as the outlier. >> >> On the open points in (1): I'm still uneasy about living with the >> snappy/lz4_raw uncompressed-size gap, and about tying is_compressed to "the >> page containing the self-reference" when the blob itself is out-of-band. >> Those feel like more reasons to prefer a small explicit layout over a >> growing set of inheritance rules. But if that's what the majority wants to >> go with I don't have a problem with it. >> >> >> On Wed, Jul 29, 2026 at 6:19 PM Alkis Evlogimenos via dev < >> [email protected]> wrote: >> >>> This was discussed at the Parquet Sync tonight. >>> >>> There is general agreement that compression is valuable for >>> self-referenced >>> data like large text blobs and that external references should remain >>> outside of Parquet's compression model. There is also agreement that we >>> need to address this before releasing the format as it renders >>> self-references impractical for text data. We should resolve this before >>> releasing the change to the format, while it can be changed cheaply. >>> >>> >>> The two positions are: >>> >>> 1. inherited compression >>> >>> The Parquet writer owns the bytes written to the file and applies the >>> same >>> compression decision whether a value is stored inline or as a >>> self-reference. >>> >>> This avoids adding per-value compression metadata, but the specification >>> must still define: >>> a. how reader obtains uncompressed size: all compressors except snappy >>> and >>> lz4raw provide this natively -> I suggest we can live with that gap. >>> b. how the compression decision is associated with a self-reference when >>> is_compressed is set to false in Data Page V2 -> I suggest we inherit the >>> is_compressed from the page containing the self-reference. This provides >>> flexibility to the writer to adjust the compression decision mid-stream. >>> c. what happens when the FILE schema does not contain an inline field -> >>> I >>> suggest that it defaults to RAW. >>> >>> 2. explicit compression metadata/framing >>> >>> Each self-reference explicitly records its compression information, >>> either >>> through additional fields or through a framing format. This is >>> self-describing and allows compression to vary between values, but adds >>> metadata and format complexity compared with the inheritance rule >>> proposed >>> in [1]. >>> >>> >>> Russel does this capture the discussion accurately? >>> >>> My preference remains inheritance. It makes the ownership model stronger: >>> the Parquet writer owns both the inline and self-reference >>> representations >>> and avoids exposing a second compression policy to users. I agree that PR >>> needs some more refinement, I will update it shortly. >>> >>> >>> >>> > Let's not rush last-minute additions after the vote ended? >>> >>> Compression inheritance was part of the proposal that was voted on and >>> was >>> removed at the last minute [2]. Regardless of how we think about that, I >>> agree we should not rush an underspecified change. At the same time we >>> should focus on resolving compressibility of text data before the first >>> release of FILE because it will be substantially easier to do so compared >>> to changing it afterward. >>> >>> [1] https://github.com/apache/parquet-format/pull/603 >>> [2] https://lists.apache.org/thread/qmx5vxg8y76xxx90cqlcfvrj7d25ps6s >>> >>> Cheers, >>> >>> On Wed, Jul 29, 2026 at 10:42 PM Antoine Pitrou <[email protected]> >>> wrote: >>> >>> > >>> > I agree with Russell here. Let's not rush last-minute additions after >>> > the vote ended? >>> > >>> > Regards >>> > >>> > Antoine. >>> > >>> > >>> > Le 29/07/2026 à 19:17, Russell Spitzer a écrit : >>> > >> >>> > >> This is equivalent to any data written in parquet and compressed as >>> a >>> > >> page, so I don't see the issue there. >>> > > >>> > > >>> > > This is the problem I have. We haven't defined how the Parquet writes >>> > these >>> > > bytes, so assuming the self reference can be treated as a page seems >>> > > undefined to me. >>> > > >>> > > On Wed, Jul 29, 2026 at 12:12 PM Alkis Evlogimenos via dev < >>> > > [email protected]> wrote: >>> > > >>> > >> I opened a PR for this change here: >>> > >> https://github.com/apache/parquet-format/pull/603 >>> > >> >>> > >> On Wed, Jul 29, 2026 at 7:49 PM Daniel Weeks <[email protected]> >>> wrote: >>> > >> >>> > >>> Russell, I'm not sure I follow your points here. With inline, the >>> data >>> > >> is >>> > >>> compressed using the column compression defined in the writer. >>> This is >>> > >>> equivalent to any data written in parquet and compressed as a >>> page, so >>> > I >>> > >>> don't see the issue there. >>> > >>> >>> > >>> The self-ref/in-file offset+size would represent the compressed >>> size >>> > >> (when >>> > >>> compressed). Having the uncompressed size is nice, but >>> technically not >>> > >>> guaranteed to be accurate because uncompressed sizes rely heavily >>> on >>> > the >>> > >>> memory layout which can vary by language/implementation. >>> > >>> >>> > >>> I think this approach largely aligns with how parquet handles data >>> > >> managed >>> > >>> by the writer. >>> > >>> >>> > >>> -Dan >>> > >>> >>> > >>> On Wed, Jul 29, 2026 at 9:37 AM Russell Spitzer < >>> > >> [email protected] >>> > >>>> >>> > >>> wrote: >>> > >>> >>> > >>>> I'm a little worried about adding this in, >>> > >>>> >>> > >>>> Self-references are already specified as [offset, offset+size) >>> ranges: >>> > >>> the >>> > >>>> resolved bytes are that range, and no compression transform is >>> > defined. >>> > >>>> Adding compression isn't a one-line codec inheritance rule. If we >>> > reuse >>> > >>>> Parquet's CompressionCodec model, readers also need framing (at >>> least >>> > >>>> uncompressed length). "The inline column chunk's >>> CompressionCodec" is >>> > >>>> underspecified too: a FILE group may omit inline, codecs are per >>> > column >>> > >>>> chunk / row group, and v2 can leave data uncompressed under the >>> same >>> > >>> chunk >>> > >>>> codec via is_compressed. >>> > >>>> >>> > >>>> Encryption is a good example of the same gap. The merged text says >>> > >>> self-ref >>> > >>>> files must not use modular encryption, but we never really worked >>> > >> through >>> > >>>> how those ranges would interact with features that need >>> page/module >>> > >>>> structure. We shouldn't bolt compression onto the same ranges >>> without >>> > >>>> defining how compressed self-ref bytes are laid out. >>> > >>>> >>> > >>>> IMHO, we should first strictly define how these out-of-band bytes >>> are >>> > >>> laid >>> > >>>> out and framed, then we can make more concrete decisions about >>> > >> inheriting >>> > >>>> codecs, encryption, and so on. >>> > >>>> >>> > >>>> On Wed, Jul 29, 2026 at 9:11 AM Alkis Evlogimenos via dev < >>> > >>>> [email protected]> wrote: >>> > >>>> >>> > >>>>> Hello, >>> > >>>>> >>> > >>>>> Now that FILE [1] is merged I'd like to reopen one point from the >>> > >>>> original >>> > >>>>> proposal that got dropped before merge: the bytes of a >>> self-reference >>> > >>>>> should use the same CompressionCodec as the column's inline >>> field. >>> > >>>> Removing >>> > >>>>> it was a mistake, and it's cheap to fix while FILE hasn't >>> shipped in >>> > >> a >>> > >>>>> release. >>> > >>>>> >>> > >>>>> As merged, a self-reference can only be stored uncompressed. >>> That's >>> > >> ok >>> > >>>> for >>> > >>>>> data like images and video, but it makes self-references useless >>> for >>> > >>> text >>> > >>>>> blobs (html, json, logs). Values too big to be inline yet small >>> > >> enough >>> > >>> to >>> > >>>>> want as self-references are exactly where PLAIN blows up the >>> storage >>> > >>>> cost. >>> > >>>>> A self-reference is the Parquet writer's decision to store a >>> large >>> > >>> value >>> > >>>>> out-of-band, so it should be compressed consistently with the >>> inline >>> > >>>> values >>> > >>>>> it came from. >>> > >>>>> >>> > >>>>> To the objections from the PR: >>> > >>>>> >>> > >>>>> 1. Apply it uniformly to external refs too. >>> > >>>>> >>> > >>>>> The asymmetry is the point. An external s3://… can be referenced >>> by >>> > >>> many >>> > >>>>> files and systems, including ones that know nothing of Parquet; >>> its >>> > >>>>> encoding is decided above Parquet, sometimes outside any engine. >>> A >>> > >>>>> self-reference lives inside Parquet and is written by the Parquet >>> > >>> writer. >>> > >>>>> Parquet owns those bytes, so Parquet compresses them. >>> > >>>>> >>> > >>>>> 2. Let the engine own it. >>> > >>>>> >>> > >>>>> The engine already owns the blob's own compression via >>> content_type >>> > >> and >>> > >>>> can >>> > >>>>> pass the writer pre-compressed bytes. The inline codec is a >>> different >>> > >>>>> thing: the storage compression Parquet applies to the column. A >>> > >>>>> self-reference is the same bytes spilled out-of-band, so it >>> belongs >>> > >> to >>> > >>>> the >>> > >>>>> same storage and codec. >>> > >>>>> >>> > >>>>> 3. Force-compressing images wastes CPU. >>> > >>>>> >>> > >>>>> It doesn't. The writer picks the codec per column chunk (and per >>> page >>> > >>> in >>> > >>>>> v2), same as it already does for inline. Inheritance just >>> propagates >>> > >>> that >>> > >>>>> choice. A column chunk written uncompressed stays uncompressed. >>> > >>>>> >>> > >>>>> 4. Compaction would force decompress/recompress. >>> > >>>>> >>> > >>>>> Only an issue for external refs. Compaction rewrites the whole >>> file, >>> > >> so >>> > >>>>> self-referenced bytes re-encode in the same pass, exactly like >>> inline >>> > >>>>> values. >>> > >>>>> >>> > >>>>> Proposed wording: >>> > >>>>> >>> > >>>>>> The bytes referenced by a self-reference (a FILE with no uri) >>> are >>> > >>>>> compressed with the CompressionCodec of the inline column chunk's >>> > >>>>> ColumnMetadata. This does not apply to external references. >>> > >>>>> >>> > >>>>> Cheers, >>> > >>>>> >>> > >>>>> [1] https://github.com/apache/parquet-format/pull/585 >>> > >>>>> >>> > >>>> >>> > >>> >>> > >> >>> > > >>> > >>> > >>> > >>> >>
