I agree with Russell here. Let's not rush last-minute additions after the vote ended?

Regards

Antoine.


Le 29/07/2026 à 19:17, Russell Spitzer a écrit :

This is equivalent to any data written in parquet and compressed as a
page, so I don't see the issue there.


This is the problem I have. We haven't defined how the Parquet writes these
bytes, so assuming the self reference can be treated as a page seems
undefined to me.

On Wed, Jul 29, 2026 at 12:12 PM Alkis Evlogimenos via dev <
[email protected]> wrote:

I opened a PR for this change here:
https://github.com/apache/parquet-format/pull/603

On Wed, Jul 29, 2026 at 7:49 PM Daniel Weeks <[email protected]> wrote:

Russell, I'm not sure I follow your points here.  With inline, the data
is
compressed using the column compression defined in the writer.  This is
equivalent to any data written in parquet and compressed as a page, so I
don't see the issue there.

The self-ref/in-file offset+size would represent the compressed size
(when
compressed).  Having the uncompressed size is nice, but technically not
guaranteed to be accurate because uncompressed sizes rely heavily on the
memory layout which can vary by language/implementation.

I think this approach largely aligns with how parquet handles data
managed
by the writer.

-Dan

On Wed, Jul 29, 2026 at 9:37 AM Russell Spitzer <
[email protected]

wrote:

I'm a little worried about adding this in,

Self-references are already specified as [offset, offset+size) ranges:
the
resolved bytes are that range, and no compression transform is defined.
Adding compression isn't a one-line codec inheritance rule. If we reuse
Parquet's CompressionCodec model, readers also need framing (at least
uncompressed length). "The inline column chunk's CompressionCodec" is
underspecified too: a FILE group may omit inline, codecs are per column
chunk / row group, and v2 can leave data uncompressed under the same
chunk
codec via is_compressed.

Encryption is a good example of the same gap. The merged text says
self-ref
files must not use modular encryption, but we never really worked
through
how those ranges would interact with features that need page/module
structure. We shouldn't bolt compression onto the same ranges without
defining how compressed self-ref bytes are laid out.

IMHO, we should first strictly define how these out-of-band bytes are
laid
out and framed, then we can make more concrete decisions about
inheriting
codecs, encryption, and so on.

On Wed, Jul 29, 2026 at 9:11 AM Alkis Evlogimenos via dev <
[email protected]> wrote:

Hello,

Now that FILE [1] is merged I'd like to reopen one point from the
original
proposal that got dropped before merge: the bytes of a self-reference
should use the same CompressionCodec as the column's inline field.
Removing
it was a mistake, and it's cheap to fix while FILE hasn't shipped in
a
release.

As merged, a self-reference can only be stored uncompressed. That's
ok
for
data like images and video, but it makes self-references useless for
text
blobs (html, json, logs). Values too big to be inline yet small
enough
to
want as self-references are exactly where PLAIN blows up the storage
cost.
A self-reference is the Parquet writer's decision to store a large
value
out-of-band, so it should be compressed consistently with the inline
values
it came from.

To the objections from the PR:

1. Apply it uniformly to external refs too.

The asymmetry is the point. An external s3://… can be referenced by
many
files and systems, including ones that know nothing of Parquet; its
encoding is decided above Parquet, sometimes outside any engine. A
self-reference lives inside Parquet and is written by the Parquet
writer.
Parquet owns those bytes, so Parquet compresses them.

2. Let the engine own it.

The engine already owns the blob's own compression via content_type
and
can
pass the writer pre-compressed bytes. The inline codec is a different
thing: the storage compression Parquet applies to the column. A
self-reference is the same bytes spilled out-of-band, so it belongs
to
the
same storage and codec.

3. Force-compressing images wastes CPU.

It doesn't. The writer picks the codec per column chunk (and per page
in
v2), same as it already does for inline. Inheritance just propagates
that
choice. A column chunk written uncompressed stays uncompressed.

4. Compaction would force decompress/recompress.

Only an issue for external refs. Compaction rewrites the whole file,
so
self-referenced bytes re-encode in the same pass, exactly like inline
values.

Proposed wording:

The bytes referenced by a self-reference (a FILE with no uri) are
compressed with the CompressionCodec of the inline column chunk's
ColumnMetadata. This does not apply to external references.

Cheers,

[1] https://github.com/apache/parquet-format/pull/585







Reply via email to