>
> This is equivalent to any data written in parquet and compressed as a
> page, so I don't see the issue there.


This is the problem I have. We haven't defined how the Parquet writes these
bytes, so assuming the self reference can be treated as a page seems
undefined to me.

On Wed, Jul 29, 2026 at 12:12 PM Alkis Evlogimenos via dev <
[email protected]> wrote:

> I opened a PR for this change here:
> https://github.com/apache/parquet-format/pull/603
>
> On Wed, Jul 29, 2026 at 7:49 PM Daniel Weeks <[email protected]> wrote:
>
> > Russell, I'm not sure I follow your points here.  With inline, the data
> is
> > compressed using the column compression defined in the writer.  This is
> > equivalent to any data written in parquet and compressed as a page, so I
> > don't see the issue there.
> >
> > The self-ref/in-file offset+size would represent the compressed size
> (when
> > compressed).  Having the uncompressed size is nice, but technically not
> > guaranteed to be accurate because uncompressed sizes rely heavily on the
> > memory layout which can vary by language/implementation.
> >
> > I think this approach largely aligns with how parquet handles data
> managed
> > by the writer.
> >
> > -Dan
> >
> > On Wed, Jul 29, 2026 at 9:37 AM Russell Spitzer <
> [email protected]
> > >
> > wrote:
> >
> > > I'm a little worried about adding this in,
> > >
> > > Self-references are already specified as [offset, offset+size) ranges:
> > the
> > > resolved bytes are that range, and no compression transform is defined.
> > > Adding compression isn't a one-line codec inheritance rule. If we reuse
> > > Parquet's CompressionCodec model, readers also need framing (at least
> > > uncompressed length). "The inline column chunk's CompressionCodec" is
> > > underspecified too: a FILE group may omit inline, codecs are per column
> > > chunk / row group, and v2 can leave data uncompressed under the same
> > chunk
> > > codec via is_compressed.
> > >
> > > Encryption is a good example of the same gap. The merged text says
> > self-ref
> > > files must not use modular encryption, but we never really worked
> through
> > > how those ranges would interact with features that need page/module
> > > structure. We shouldn't bolt compression onto the same ranges without
> > > defining how compressed self-ref bytes are laid out.
> > >
> > > IMHO, we should first strictly define how these out-of-band bytes are
> > laid
> > > out and framed, then we can make more concrete decisions about
> inheriting
> > > codecs, encryption, and so on.
> > >
> > > On Wed, Jul 29, 2026 at 9:11 AM Alkis Evlogimenos via dev <
> > > [email protected]> wrote:
> > >
> > > > Hello,
> > > >
> > > > Now that FILE [1] is merged I'd like to reopen one point from the
> > > original
> > > > proposal that got dropped before merge: the bytes of a self-reference
> > > > should use the same CompressionCodec as the column's inline field.
> > > Removing
> > > > it was a mistake, and it's cheap to fix while FILE hasn't shipped in
> a
> > > > release.
> > > >
> > > > As merged, a self-reference can only be stored uncompressed. That's
> ok
> > > for
> > > > data like images and video, but it makes self-references useless for
> > text
> > > > blobs (html, json, logs). Values too big to be inline yet small
> enough
> > to
> > > > want as self-references are exactly where PLAIN blows up the storage
> > > cost.
> > > > A self-reference is the Parquet writer's decision to store a large
> > value
> > > > out-of-band, so it should be compressed consistently with the inline
> > > values
> > > > it came from.
> > > >
> > > > To the objections from the PR:
> > > >
> > > > 1. Apply it uniformly to external refs too.
> > > >
> > > > The asymmetry is the point. An external s3://… can be referenced by
> > many
> > > > files and systems, including ones that know nothing of Parquet; its
> > > > encoding is decided above Parquet, sometimes outside any engine. A
> > > > self-reference lives inside Parquet and is written by the Parquet
> > writer.
> > > > Parquet owns those bytes, so Parquet compresses them.
> > > >
> > > > 2. Let the engine own it.
> > > >
> > > > The engine already owns the blob's own compression via content_type
> and
> > > can
> > > > pass the writer pre-compressed bytes. The inline codec is a different
> > > > thing: the storage compression Parquet applies to the column. A
> > > > self-reference is the same bytes spilled out-of-band, so it belongs
> to
> > > the
> > > > same storage and codec.
> > > >
> > > > 3. Force-compressing images wastes CPU.
> > > >
> > > > It doesn't. The writer picks the codec per column chunk (and per page
> > in
> > > > v2), same as it already does for inline. Inheritance just propagates
> > that
> > > > choice. A column chunk written uncompressed stays uncompressed.
> > > >
> > > > 4. Compaction would force decompress/recompress.
> > > >
> > > > Only an issue for external refs. Compaction rewrites the whole file,
> so
> > > > self-referenced bytes re-encode in the same pass, exactly like inline
> > > > values.
> > > >
> > > > Proposed wording:
> > > >
> > > > > The bytes referenced by a self-reference (a FILE with no uri) are
> > > > compressed with the CompressionCodec of the inline column chunk's
> > > > ColumnMetadata. This does not apply to external references.
> > > >
> > > > Cheers,
> > > >
> > > > [1] https://github.com/apache/parquet-format/pull/585
> > > >
> > >
> >
>

Reply via email to