I updated https://github.com/apache/parquet-format/pull/603 to remove
self-references. There is a discussion on the PR right now.

One thing we discovered from the reviews was that the current spec as
merged is underspecified on if we allow `inline` and `uri+offset+len`. I
could see some value of `inline` being a cached copy of the referred bytes
so I am tweaking the wording to allow this and say that reader can choose
to read any representation and expect the bytes to be the same.

Once the dust settles on the PR I will start a VOTE with the changes before
merging.


On Wed, Aug 19, 2026 at 7:00 AM Russell Spitzer <[email protected]>
wrote:

> That sounds a lot more sustainable long term, so I'm excited to see the
> proposal.
>
> On Tue, Aug 18, 2026 at 10:22 PM Alkis Evlogimenos <
> [email protected]> wrote:
>
>> I am withdrawing this. PR 603 [1] now removes self-references from FILE
>> instead of specifying compression and encryption for them.
>>
>> Through extensive discussion with folks we think that a non-contiguous
>> page mechanism belongs to the non-contiguous page proposal and not as a
>> special way to do FILE. This means we should disallow self-references in
>> FILE and let the non-contiguous pages proposal to allow storing outside
>> values of the `inline` column elsewhere in the file. If those values are in
>> pages they naturally get al the fields: `is_compressed`, `encryption`,
>> `uncompressed_page_size` etc. As such non-contiguous pages composes with
>> FILE instead of competing with it.
>>
>> We will send the non-contiguous pages proposal in a separate thread.
>>
>> Cheers,
>>
>> [1] https://github.com/apache/parquet-format/pull/603
>>
>>
>> On Thu, Jul 30, 2026 at 9:55 AM Alkis Evlogimenos <
>> [email protected]> wrote:
>>
>>> Great!
>>>
>>> I have updated the PR [1] with precise semantics and added encryption
>>> rules.
>>>
>>> The remaining question is the uncompressed size. I did a deeper analysis
>>> of the codecs. Snappy includes it in its compressed representation, but
>>> several other Parquet codecs either omit it or do not guarantee it. The
>>> current proposal requires readers to support dynamically sized
>>> decompression output. I think that's an acceptable tradeoff but I am open
>>> to alternatives if there is strong need to know the uncompressed size
>>> before decompression.
>>>
>>> [1] https://github.com/apache/parquet-format/pull/603
>>>
>>> On Thu, Jul 30, 2026 at 3:45 AM Russell Spitzer <
>>> [email protected]> wrote:
>>>
>>>> I think that's pretty accurate. We should probably also establish
>>>> encryption rules concurrently with any compression rules we add.
>>>>
>>>> If I can summarize the pro argument:
>>>>
>>>> "A Parquet writer is already writing these bytes; had they been inlined
>>>> they would have had compression applied. Therefore, if a writer is writing
>>>> a blob into non-columnar space, it should use the same compression as it
>>>> would have if it were writing into an inline page."
>>>>
>>>> My initial hesitation was that I thought the consumer of this object
>>>> would want to interact with the blob using only the offset. Daniel's
>>>> perspective is that the Parquet reader will always be an interlocutor for
>>>> getting these bytes, so we can rely on inheritance from Parquet internals
>>>> to produce the actual byte stream. That clarifies the model, and it's also
>>>> why I'm more in favor of an explicit definition here. Inheritance means we
>>>> cannot use the offset on its own, so we essentially have three modes of
>>>> reading:
>>>>
>>>> Inline - The Parquet reader returns the bytes in the type object.
>>>> Self-reference - The reader could return offset/size, but that may be
>>>> meaningless or miss compression information. If an engine wants the logical
>>>> bytes, it needs a byte stream from the Parquet reader and can't just copy
>>>> the range as-is from the file.
>>>> External reference - The reader returns a path; it's up to the consumer
>>>> to know what to do with it.
>>>>
>>>> I thought we were basically doing self reference and external reference
>>>> the same way: whether there is compression or not is a function of the
>>>> actual file contents, not of the fact that the range lives inside a Parquet
>>>> file. I can see the benefit of inheriting, but in that case I think we
>>>> should treat self-ref bytes like other writer-owned storage (compression,
>>>> encryption, and anything similar we add later), and not just add only on
>>>> codec inheritance. I think the pro inheritance case is basically saying we
>>>> should have inline and self-reference behave the same with external
>>>> reference as the outlier.
>>>>
>>>> On the open points in (1): I'm still uneasy about living with the
>>>> snappy/lz4_raw uncompressed-size gap, and about tying is_compressed to "the
>>>> page containing the self-reference" when the blob itself is out-of-band.
>>>> Those feel like more reasons to prefer a small explicit layout over a
>>>> growing set of inheritance rules. But if that's what the majority wants to
>>>> go with I don't have a problem with it.
>>>>
>>>>
>>>> On Wed, Jul 29, 2026 at 6:19 PM Alkis Evlogimenos via dev <
>>>> [email protected]> wrote:
>>>>
>>>>> This was discussed at the Parquet Sync tonight.
>>>>>
>>>>> There is general agreement that compression is valuable for
>>>>> self-referenced
>>>>> data like large text blobs and that external references should remain
>>>>> outside of Parquet's compression model. There is also agreement that we
>>>>> need to address this before releasing the format as it renders
>>>>> self-references impractical for text data. We should resolve this
>>>>> before
>>>>> releasing the change to the format, while it can be changed cheaply.
>>>>>
>>>>>
>>>>> The two positions are:
>>>>>
>>>>> 1. inherited compression
>>>>>
>>>>> The Parquet writer owns the bytes written to the file and applies the
>>>>> same
>>>>> compression decision whether a value is stored inline or as a
>>>>> self-reference.
>>>>>
>>>>> This avoids adding per-value compression metadata, but the
>>>>> specification
>>>>> must still define:
>>>>> a. how reader obtains uncompressed size: all compressors except snappy
>>>>> and
>>>>> lz4raw provide this natively -> I suggest we can live with that gap.
>>>>> b. how the compression decision is associated with a self-reference
>>>>> when
>>>>> is_compressed is set to false in Data Page V2 -> I suggest we inherit
>>>>> the
>>>>> is_compressed from the page containing the self-reference. This
>>>>> provides
>>>>> flexibility to the writer to adjust the compression decision
>>>>> mid-stream.
>>>>> c. what happens when the FILE schema does not contain an inline field
>>>>> -> I
>>>>> suggest that it defaults to RAW.
>>>>>
>>>>> 2. explicit compression metadata/framing
>>>>>
>>>>> Each self-reference explicitly records its compression information,
>>>>> either
>>>>> through additional fields or through a framing format. This is
>>>>> self-describing and allows compression to vary between values, but adds
>>>>> metadata and format complexity compared with the inheritance rule
>>>>> proposed
>>>>> in [1].
>>>>>
>>>>>
>>>>> Russel does this capture the discussion accurately?
>>>>>
>>>>> My preference remains inheritance. It makes the ownership model
>>>>> stronger:
>>>>> the Parquet writer owns both the inline and self-reference
>>>>> representations
>>>>> and avoids exposing a second compression policy to users. I agree that
>>>>> PR
>>>>> needs some more refinement, I will update it shortly.
>>>>>
>>>>>
>>>>>
>>>>> > Let's not rush last-minute additions after the vote ended?
>>>>>
>>>>> Compression inheritance was part of the proposal that was voted on and
>>>>> was
>>>>> removed at the last minute [2]. Regardless of how we think about that,
>>>>> I
>>>>> agree we should not rush an underspecified change. At the same time we
>>>>> should focus on resolving compressibility of text data before the first
>>>>> release of FILE because it will be substantially easier to do so
>>>>> compared
>>>>> to changing it afterward.
>>>>>
>>>>> [1] https://github.com/apache/parquet-format/pull/603
>>>>> [2] https://lists.apache.org/thread/qmx5vxg8y76xxx90cqlcfvrj7d25ps6s
>>>>>
>>>>> Cheers,
>>>>>
>>>>> On Wed, Jul 29, 2026 at 10:42 PM Antoine Pitrou <[email protected]>
>>>>> wrote:
>>>>>
>>>>> >
>>>>> > I agree with Russell here. Let's not rush last-minute additions after
>>>>> > the vote ended?
>>>>> >
>>>>> > Regards
>>>>> >
>>>>> > Antoine.
>>>>> >
>>>>> >
>>>>> > Le 29/07/2026 à 19:17, Russell Spitzer a écrit :
>>>>> > >>
>>>>> > >> This is equivalent to any data written in parquet and compressed
>>>>> as a
>>>>> > >> page, so I don't see the issue there.
>>>>> > >
>>>>> > >
>>>>> > > This is the problem I have. We haven't defined how the Parquet
>>>>> writes
>>>>> > these
>>>>> > > bytes, so assuming the self reference can be treated as a page
>>>>> seems
>>>>> > > undefined to me.
>>>>> > >
>>>>> > > On Wed, Jul 29, 2026 at 12:12 PM Alkis Evlogimenos via dev <
>>>>> > > [email protected]> wrote:
>>>>> > >
>>>>> > >> I opened a PR for this change here:
>>>>> > >> https://github.com/apache/parquet-format/pull/603
>>>>> > >>
>>>>> > >> On Wed, Jul 29, 2026 at 7:49 PM Daniel Weeks <[email protected]>
>>>>> wrote:
>>>>> > >>
>>>>> > >>> Russell, I'm not sure I follow your points here.  With inline,
>>>>> the data
>>>>> > >> is
>>>>> > >>> compressed using the column compression defined in the writer.
>>>>> This is
>>>>> > >>> equivalent to any data written in parquet and compressed as a
>>>>> page, so
>>>>> > I
>>>>> > >>> don't see the issue there.
>>>>> > >>>
>>>>> > >>> The self-ref/in-file offset+size would represent the compressed
>>>>> size
>>>>> > >> (when
>>>>> > >>> compressed).  Having the uncompressed size is nice, but
>>>>> technically not
>>>>> > >>> guaranteed to be accurate because uncompressed sizes rely
>>>>> heavily on
>>>>> > the
>>>>> > >>> memory layout which can vary by language/implementation.
>>>>> > >>>
>>>>> > >>> I think this approach largely aligns with how parquet handles
>>>>> data
>>>>> > >> managed
>>>>> > >>> by the writer.
>>>>> > >>>
>>>>> > >>> -Dan
>>>>> > >>>
>>>>> > >>> On Wed, Jul 29, 2026 at 9:37 AM Russell Spitzer <
>>>>> > >> [email protected]
>>>>> > >>>>
>>>>> > >>> wrote:
>>>>> > >>>
>>>>> > >>>> I'm a little worried about adding this in,
>>>>> > >>>>
>>>>> > >>>> Self-references are already specified as [offset, offset+size)
>>>>> ranges:
>>>>> > >>> the
>>>>> > >>>> resolved bytes are that range, and no compression transform is
>>>>> > defined.
>>>>> > >>>> Adding compression isn't a one-line codec inheritance rule. If
>>>>> we
>>>>> > reuse
>>>>> > >>>> Parquet's CompressionCodec model, readers also need framing (at
>>>>> least
>>>>> > >>>> uncompressed length). "The inline column chunk's
>>>>> CompressionCodec" is
>>>>> > >>>> underspecified too: a FILE group may omit inline, codecs are per
>>>>> > column
>>>>> > >>>> chunk / row group, and v2 can leave data uncompressed under the
>>>>> same
>>>>> > >>> chunk
>>>>> > >>>> codec via is_compressed.
>>>>> > >>>>
>>>>> > >>>> Encryption is a good example of the same gap. The merged text
>>>>> says
>>>>> > >>> self-ref
>>>>> > >>>> files must not use modular encryption, but we never really
>>>>> worked
>>>>> > >> through
>>>>> > >>>> how those ranges would interact with features that need
>>>>> page/module
>>>>> > >>>> structure. We shouldn't bolt compression onto the same ranges
>>>>> without
>>>>> > >>>> defining how compressed self-ref bytes are laid out.
>>>>> > >>>>
>>>>> > >>>> IMHO, we should first strictly define how these out-of-band
>>>>> bytes are
>>>>> > >>> laid
>>>>> > >>>> out and framed, then we can make more concrete decisions about
>>>>> > >> inheriting
>>>>> > >>>> codecs, encryption, and so on.
>>>>> > >>>>
>>>>> > >>>> On Wed, Jul 29, 2026 at 9:11 AM Alkis Evlogimenos via dev <
>>>>> > >>>> [email protected]> wrote:
>>>>> > >>>>
>>>>> > >>>>> Hello,
>>>>> > >>>>>
>>>>> > >>>>> Now that FILE [1] is merged I'd like to reopen one point from
>>>>> the
>>>>> > >>>> original
>>>>> > >>>>> proposal that got dropped before merge: the bytes of a
>>>>> self-reference
>>>>> > >>>>> should use the same CompressionCodec as the column's inline
>>>>> field.
>>>>> > >>>> Removing
>>>>> > >>>>> it was a mistake, and it's cheap to fix while FILE hasn't
>>>>> shipped in
>>>>> > >> a
>>>>> > >>>>> release.
>>>>> > >>>>>
>>>>> > >>>>> As merged, a self-reference can only be stored uncompressed.
>>>>> That's
>>>>> > >> ok
>>>>> > >>>> for
>>>>> > >>>>> data like images and video, but it makes self-references
>>>>> useless for
>>>>> > >>> text
>>>>> > >>>>> blobs (html, json, logs). Values too big to be inline yet small
>>>>> > >> enough
>>>>> > >>> to
>>>>> > >>>>> want as self-references are exactly where PLAIN blows up the
>>>>> storage
>>>>> > >>>> cost.
>>>>> > >>>>> A self-reference is the Parquet writer's decision to store a
>>>>> large
>>>>> > >>> value
>>>>> > >>>>> out-of-band, so it should be compressed consistently with the
>>>>> inline
>>>>> > >>>> values
>>>>> > >>>>> it came from.
>>>>> > >>>>>
>>>>> > >>>>> To the objections from the PR:
>>>>> > >>>>>
>>>>> > >>>>> 1. Apply it uniformly to external refs too.
>>>>> > >>>>>
>>>>> > >>>>> The asymmetry is the point. An external s3://… can be
>>>>> referenced by
>>>>> > >>> many
>>>>> > >>>>> files and systems, including ones that know nothing of
>>>>> Parquet; its
>>>>> > >>>>> encoding is decided above Parquet, sometimes outside any
>>>>> engine. A
>>>>> > >>>>> self-reference lives inside Parquet and is written by the
>>>>> Parquet
>>>>> > >>> writer.
>>>>> > >>>>> Parquet owns those bytes, so Parquet compresses them.
>>>>> > >>>>>
>>>>> > >>>>> 2. Let the engine own it.
>>>>> > >>>>>
>>>>> > >>>>> The engine already owns the blob's own compression via
>>>>> content_type
>>>>> > >> and
>>>>> > >>>> can
>>>>> > >>>>> pass the writer pre-compressed bytes. The inline codec is a
>>>>> different
>>>>> > >>>>> thing: the storage compression Parquet applies to the column. A
>>>>> > >>>>> self-reference is the same bytes spilled out-of-band, so it
>>>>> belongs
>>>>> > >> to
>>>>> > >>>> the
>>>>> > >>>>> same storage and codec.
>>>>> > >>>>>
>>>>> > >>>>> 3. Force-compressing images wastes CPU.
>>>>> > >>>>>
>>>>> > >>>>> It doesn't. The writer picks the codec per column chunk (and
>>>>> per page
>>>>> > >>> in
>>>>> > >>>>> v2), same as it already does for inline. Inheritance just
>>>>> propagates
>>>>> > >>> that
>>>>> > >>>>> choice. A column chunk written uncompressed stays uncompressed.
>>>>> > >>>>>
>>>>> > >>>>> 4. Compaction would force decompress/recompress.
>>>>> > >>>>>
>>>>> > >>>>> Only an issue for external refs. Compaction rewrites the whole
>>>>> file,
>>>>> > >> so
>>>>> > >>>>> self-referenced bytes re-encode in the same pass, exactly like
>>>>> inline
>>>>> > >>>>> values.
>>>>> > >>>>>
>>>>> > >>>>> Proposed wording:
>>>>> > >>>>>
>>>>> > >>>>>> The bytes referenced by a self-reference (a FILE with no uri)
>>>>> are
>>>>> > >>>>> compressed with the CompressionCodec of the inline column
>>>>> chunk's
>>>>> > >>>>> ColumnMetadata. This does not apply to external references.
>>>>> > >>>>>
>>>>> > >>>>> Cheers,
>>>>> > >>>>>
>>>>> > >>>>> [1] https://github.com/apache/parquet-format/pull/585
>>>>> > >>>>>
>>>>> > >>>>
>>>>> > >>>
>>>>> > >>
>>>>> > >
>>>>> >
>>>>> >
>>>>> >
>>>>>
>>>>

Reply via email to