I am withdrawing this. PR 603 [1] now removes self-references from FILE
instead of specifying compression and encryption for them.

Through extensive discussion with folks we think that a non-contiguous page
mechanism belongs to the non-contiguous page proposal and not as a special
way to do FILE. This means we should disallow self-references in FILE and
let the non-contiguous pages proposal to allow storing outside values of
the `inline` column elsewhere in the file. If those values are in pages
they naturally get al the fields: `is_compressed`, `encryption`,
`uncompressed_page_size` etc. As such non-contiguous pages composes with
FILE instead of competing with it.

We will send the non-contiguous pages proposal in a separate thread.

Cheers,

[1] https://github.com/apache/parquet-format/pull/603


On Thu, Jul 30, 2026 at 9:55 AM Alkis Evlogimenos <
[email protected]> wrote:

> Great!
>
> I have updated the PR [1] with precise semantics and added encryption
> rules.
>
> The remaining question is the uncompressed size. I did a deeper analysis
> of the codecs. Snappy includes it in its compressed representation, but
> several other Parquet codecs either omit it or do not guarantee it. The
> current proposal requires readers to support dynamically sized
> decompression output. I think that's an acceptable tradeoff but I am open
> to alternatives if there is strong need to know the uncompressed size
> before decompression.
>
> [1] https://github.com/apache/parquet-format/pull/603
>
> On Thu, Jul 30, 2026 at 3:45 AM Russell Spitzer <[email protected]>
> wrote:
>
>> I think that's pretty accurate. We should probably also establish
>> encryption rules concurrently with any compression rules we add.
>>
>> If I can summarize the pro argument:
>>
>> "A Parquet writer is already writing these bytes; had they been inlined
>> they would have had compression applied. Therefore, if a writer is writing
>> a blob into non-columnar space, it should use the same compression as it
>> would have if it were writing into an inline page."
>>
>> My initial hesitation was that I thought the consumer of this object
>> would want to interact with the blob using only the offset. Daniel's
>> perspective is that the Parquet reader will always be an interlocutor for
>> getting these bytes, so we can rely on inheritance from Parquet internals
>> to produce the actual byte stream. That clarifies the model, and it's also
>> why I'm more in favor of an explicit definition here. Inheritance means we
>> cannot use the offset on its own, so we essentially have three modes of
>> reading:
>>
>> Inline - The Parquet reader returns the bytes in the type object.
>> Self-reference - The reader could return offset/size, but that may be
>> meaningless or miss compression information. If an engine wants the logical
>> bytes, it needs a byte stream from the Parquet reader and can't just copy
>> the range as-is from the file.
>> External reference - The reader returns a path; it's up to the consumer
>> to know what to do with it.
>>
>> I thought we were basically doing self reference and external reference
>> the same way: whether there is compression or not is a function of the
>> actual file contents, not of the fact that the range lives inside a Parquet
>> file. I can see the benefit of inheriting, but in that case I think we
>> should treat self-ref bytes like other writer-owned storage (compression,
>> encryption, and anything similar we add later), and not just add only on
>> codec inheritance. I think the pro inheritance case is basically saying we
>> should have inline and self-reference behave the same with external
>> reference as the outlier.
>>
>> On the open points in (1): I'm still uneasy about living with the
>> snappy/lz4_raw uncompressed-size gap, and about tying is_compressed to "the
>> page containing the self-reference" when the blob itself is out-of-band.
>> Those feel like more reasons to prefer a small explicit layout over a
>> growing set of inheritance rules. But if that's what the majority wants to
>> go with I don't have a problem with it.
>>
>>
>> On Wed, Jul 29, 2026 at 6:19 PM Alkis Evlogimenos via dev <
>> [email protected]> wrote:
>>
>>> This was discussed at the Parquet Sync tonight.
>>>
>>> There is general agreement that compression is valuable for
>>> self-referenced
>>> data like large text blobs and that external references should remain
>>> outside of Parquet's compression model. There is also agreement that we
>>> need to address this before releasing the format as it renders
>>> self-references impractical for text data. We should resolve this before
>>> releasing the change to the format, while it can be changed cheaply.
>>>
>>>
>>> The two positions are:
>>>
>>> 1. inherited compression
>>>
>>> The Parquet writer owns the bytes written to the file and applies the
>>> same
>>> compression decision whether a value is stored inline or as a
>>> self-reference.
>>>
>>> This avoids adding per-value compression metadata, but the specification
>>> must still define:
>>> a. how reader obtains uncompressed size: all compressors except snappy
>>> and
>>> lz4raw provide this natively -> I suggest we can live with that gap.
>>> b. how the compression decision is associated with a self-reference when
>>> is_compressed is set to false in Data Page V2 -> I suggest we inherit the
>>> is_compressed from the page containing the self-reference. This provides
>>> flexibility to the writer to adjust the compression decision mid-stream.
>>> c. what happens when the FILE schema does not contain an inline field ->
>>> I
>>> suggest that it defaults to RAW.
>>>
>>> 2. explicit compression metadata/framing
>>>
>>> Each self-reference explicitly records its compression information,
>>> either
>>> through additional fields or through a framing format. This is
>>> self-describing and allows compression to vary between values, but adds
>>> metadata and format complexity compared with the inheritance rule
>>> proposed
>>> in [1].
>>>
>>>
>>> Russel does this capture the discussion accurately?
>>>
>>> My preference remains inheritance. It makes the ownership model stronger:
>>> the Parquet writer owns both the inline and self-reference
>>> representations
>>> and avoids exposing a second compression policy to users. I agree that PR
>>> needs some more refinement, I will update it shortly.
>>>
>>>
>>>
>>> > Let's not rush last-minute additions after the vote ended?
>>>
>>> Compression inheritance was part of the proposal that was voted on and
>>> was
>>> removed at the last minute [2]. Regardless of how we think about that, I
>>> agree we should not rush an underspecified change. At the same time we
>>> should focus on resolving compressibility of text data before the first
>>> release of FILE because it will be substantially easier to do so compared
>>> to changing it afterward.
>>>
>>> [1] https://github.com/apache/parquet-format/pull/603
>>> [2] https://lists.apache.org/thread/qmx5vxg8y76xxx90cqlcfvrj7d25ps6s
>>>
>>> Cheers,
>>>
>>> On Wed, Jul 29, 2026 at 10:42 PM Antoine Pitrou <[email protected]>
>>> wrote:
>>>
>>> >
>>> > I agree with Russell here. Let's not rush last-minute additions after
>>> > the vote ended?
>>> >
>>> > Regards
>>> >
>>> > Antoine.
>>> >
>>> >
>>> > Le 29/07/2026 à 19:17, Russell Spitzer a écrit :
>>> > >>
>>> > >> This is equivalent to any data written in parquet and compressed as
>>> a
>>> > >> page, so I don't see the issue there.
>>> > >
>>> > >
>>> > > This is the problem I have. We haven't defined how the Parquet writes
>>> > these
>>> > > bytes, so assuming the self reference can be treated as a page seems
>>> > > undefined to me.
>>> > >
>>> > > On Wed, Jul 29, 2026 at 12:12 PM Alkis Evlogimenos via dev <
>>> > > [email protected]> wrote:
>>> > >
>>> > >> I opened a PR for this change here:
>>> > >> https://github.com/apache/parquet-format/pull/603
>>> > >>
>>> > >> On Wed, Jul 29, 2026 at 7:49 PM Daniel Weeks <[email protected]>
>>> wrote:
>>> > >>
>>> > >>> Russell, I'm not sure I follow your points here.  With inline, the
>>> data
>>> > >> is
>>> > >>> compressed using the column compression defined in the writer.
>>> This is
>>> > >>> equivalent to any data written in parquet and compressed as a
>>> page, so
>>> > I
>>> > >>> don't see the issue there.
>>> > >>>
>>> > >>> The self-ref/in-file offset+size would represent the compressed
>>> size
>>> > >> (when
>>> > >>> compressed).  Having the uncompressed size is nice, but
>>> technically not
>>> > >>> guaranteed to be accurate because uncompressed sizes rely heavily
>>> on
>>> > the
>>> > >>> memory layout which can vary by language/implementation.
>>> > >>>
>>> > >>> I think this approach largely aligns with how parquet handles data
>>> > >> managed
>>> > >>> by the writer.
>>> > >>>
>>> > >>> -Dan
>>> > >>>
>>> > >>> On Wed, Jul 29, 2026 at 9:37 AM Russell Spitzer <
>>> > >> [email protected]
>>> > >>>>
>>> > >>> wrote:
>>> > >>>
>>> > >>>> I'm a little worried about adding this in,
>>> > >>>>
>>> > >>>> Self-references are already specified as [offset, offset+size)
>>> ranges:
>>> > >>> the
>>> > >>>> resolved bytes are that range, and no compression transform is
>>> > defined.
>>> > >>>> Adding compression isn't a one-line codec inheritance rule. If we
>>> > reuse
>>> > >>>> Parquet's CompressionCodec model, readers also need framing (at
>>> least
>>> > >>>> uncompressed length). "The inline column chunk's
>>> CompressionCodec" is
>>> > >>>> underspecified too: a FILE group may omit inline, codecs are per
>>> > column
>>> > >>>> chunk / row group, and v2 can leave data uncompressed under the
>>> same
>>> > >>> chunk
>>> > >>>> codec via is_compressed.
>>> > >>>>
>>> > >>>> Encryption is a good example of the same gap. The merged text says
>>> > >>> self-ref
>>> > >>>> files must not use modular encryption, but we never really worked
>>> > >> through
>>> > >>>> how those ranges would interact with features that need
>>> page/module
>>> > >>>> structure. We shouldn't bolt compression onto the same ranges
>>> without
>>> > >>>> defining how compressed self-ref bytes are laid out.
>>> > >>>>
>>> > >>>> IMHO, we should first strictly define how these out-of-band bytes
>>> are
>>> > >>> laid
>>> > >>>> out and framed, then we can make more concrete decisions about
>>> > >> inheriting
>>> > >>>> codecs, encryption, and so on.
>>> > >>>>
>>> > >>>> On Wed, Jul 29, 2026 at 9:11 AM Alkis Evlogimenos via dev <
>>> > >>>> [email protected]> wrote:
>>> > >>>>
>>> > >>>>> Hello,
>>> > >>>>>
>>> > >>>>> Now that FILE [1] is merged I'd like to reopen one point from the
>>> > >>>> original
>>> > >>>>> proposal that got dropped before merge: the bytes of a
>>> self-reference
>>> > >>>>> should use the same CompressionCodec as the column's inline
>>> field.
>>> > >>>> Removing
>>> > >>>>> it was a mistake, and it's cheap to fix while FILE hasn't
>>> shipped in
>>> > >> a
>>> > >>>>> release.
>>> > >>>>>
>>> > >>>>> As merged, a self-reference can only be stored uncompressed.
>>> That's
>>> > >> ok
>>> > >>>> for
>>> > >>>>> data like images and video, but it makes self-references useless
>>> for
>>> > >>> text
>>> > >>>>> blobs (html, json, logs). Values too big to be inline yet small
>>> > >> enough
>>> > >>> to
>>> > >>>>> want as self-references are exactly where PLAIN blows up the
>>> storage
>>> > >>>> cost.
>>> > >>>>> A self-reference is the Parquet writer's decision to store a
>>> large
>>> > >>> value
>>> > >>>>> out-of-band, so it should be compressed consistently with the
>>> inline
>>> > >>>> values
>>> > >>>>> it came from.
>>> > >>>>>
>>> > >>>>> To the objections from the PR:
>>> > >>>>>
>>> > >>>>> 1. Apply it uniformly to external refs too.
>>> > >>>>>
>>> > >>>>> The asymmetry is the point. An external s3://… can be referenced
>>> by
>>> > >>> many
>>> > >>>>> files and systems, including ones that know nothing of Parquet;
>>> its
>>> > >>>>> encoding is decided above Parquet, sometimes outside any engine.
>>> A
>>> > >>>>> self-reference lives inside Parquet and is written by the Parquet
>>> > >>> writer.
>>> > >>>>> Parquet owns those bytes, so Parquet compresses them.
>>> > >>>>>
>>> > >>>>> 2. Let the engine own it.
>>> > >>>>>
>>> > >>>>> The engine already owns the blob's own compression via
>>> content_type
>>> > >> and
>>> > >>>> can
>>> > >>>>> pass the writer pre-compressed bytes. The inline codec is a
>>> different
>>> > >>>>> thing: the storage compression Parquet applies to the column. A
>>> > >>>>> self-reference is the same bytes spilled out-of-band, so it
>>> belongs
>>> > >> to
>>> > >>>> the
>>> > >>>>> same storage and codec.
>>> > >>>>>
>>> > >>>>> 3. Force-compressing images wastes CPU.
>>> > >>>>>
>>> > >>>>> It doesn't. The writer picks the codec per column chunk (and per
>>> page
>>> > >>> in
>>> > >>>>> v2), same as it already does for inline. Inheritance just
>>> propagates
>>> > >>> that
>>> > >>>>> choice. A column chunk written uncompressed stays uncompressed.
>>> > >>>>>
>>> > >>>>> 4. Compaction would force decompress/recompress.
>>> > >>>>>
>>> > >>>>> Only an issue for external refs. Compaction rewrites the whole
>>> file,
>>> > >> so
>>> > >>>>> self-referenced bytes re-encode in the same pass, exactly like
>>> inline
>>> > >>>>> values.
>>> > >>>>>
>>> > >>>>> Proposed wording:
>>> > >>>>>
>>> > >>>>>> The bytes referenced by a self-reference (a FILE with no uri)
>>> are
>>> > >>>>> compressed with the CompressionCodec of the inline column chunk's
>>> > >>>>> ColumnMetadata. This does not apply to external references.
>>> > >>>>>
>>> > >>>>> Cheers,
>>> > >>>>>
>>> > >>>>> [1] https://github.com/apache/parquet-format/pull/585
>>> > >>>>>
>>> > >>>>
>>> > >>>
>>> > >>
>>> > >
>>> >
>>> >
>>> >
>>>
>>

Reply via email to