Thanks for the comments!

Don't we also want to know whether the row groups align during scan
planning? If they align, we can assume that a task would read
significantly less bytes than if they don't.

On Wed, Aug 12, 2026 at 7:58 PM Anoop Johnson <[email protected]> wrote:
>
> I don't see any obvious advantage of keeping track of row group alignment in 
> the column file metadata. As Russell pointed out, the readers must read the 
> Parquet footers anyway before any decoding can happen. At that point, you can 
> derive whether the row groups are aligned by simple integer comparison. Once 
> the reader figures out the alignment, then the optimization Leonid mentioned 
> can still work?
>
> The closest precedent is where we keep track of the `split_offsets` in the 
> metadata. But that is warranted because we would like to know the offsets at 
> scan planning time before any Parquet footers are actually opened. That too 
> is an optional optimization - if the split_offsets are absent, we fall back 
> to fixed block split computation.
>
> Best,
> Anoop
>
> On Wed, Aug 12, 2026 at 9:06 AM Russell Spitzer <[email protected]> 
> wrote:
>>
>> Before adding format-specific performance optimizations like this, I think 
>> we should ensure we are solving a real problem and that it is only solvable 
>> at the metadata layer.
>>
>> In this case, can’t the reader tell at read time whether row groups are 
>> aligned by comparing footers? The reader must be given both the Base File 
>> (BF) and Column File (CF) regardless of alignment, and must open both 
>> footers before decoding. At that point it already knows whether BF and CF 
>> share the same row-group row boundaries, and can plan further I/O 
>> accordingly (1:1 chunk reads when aligned; overlapping CF row groups when 
>> not).
>>
>> Do we get a material benefit from knowing alignment ahead of time in Iceberg 
>> metadata, versus deriving it from the footers we have to read anyway?
>>
>> On Wed, Aug 12, 2026 at 10:44 AM Péter Váry <[email protected]> 
>> wrote:
>>>
>>> Hi Leonid,
>>>
>>> Thanks for the proposal and for continuing the discussion.
>>>
>>> I have a couple of questions:
>>>
>>> Do you have any suggestions for how writers could reliably produce aligned 
>>> row groups? In particular, how would an update writer obtain the base 
>>> file's row-group boundaries in practice? Would all writers be expected to 
>>> read the base file footer? Also, is there an easy way to implement this 
>>> using the current Java Parquet writer APIs? This sounds like a nice feature.
>>> Do readers actually need to know in advance that row groups are aligned? It 
>>> seems a reader could use essentially the same algorithm for both aligned 
>>> and unaligned files: seek to the nearest row group, discard any rows prior 
>>> to the desired starting position, and continue reading from there. If 
>>> page-level skipping is available, it may even avoid reading some of those 
>>> unnecessary pages. When the row group already begins at the required 
>>> offset, the discard step simply becomes a no-op. If that is the case, do we 
>>> actually need an alignment flag at all?
>>>
>>> Thanks, Peter
>>>
>>> Leonid Lygin via dev <[email protected]> ezt írta (időpont: 2026. aug. 
>>> 12., Sze, 17:06):
>>>>
>>>> Hi all!
>>>>
>>>> Following up on the Column File representation thread
>>>> <https://lists.apache.org/thread/jbh1gbrso5h6l4by9rh9poy2cjjtb8j0>, I'd 
>>>> like to
>>>> fire off a discussion about a possible optimization for Column Updates, 
>>>> where
>>>> supporting writers might decide to write out the Column File aligning all 
>>>> row
>>>> groups with the Base File, enabling supporting readers to do simple 
>>>> zero-copy
>>>> reads.
>>>>
>>>> *Context*
>>>>
>>>> The original thread settled on a dense representation (i.e. Column Files 
>>>> contain
>>>> exactly the same amount of rows as Base Files), while allowing 
>>>> unsynchronized
>>>> row-group boundaries.
>>>>
>>>> I'm proposing adding a separate metadata flag (e.g.
>>>> `ColumnFile.containsAlignedRowGroups`) which, when set, indicates that the 
>>>> row
>>>> groups contained in the associated Column File are aligned with the Base 
>>>> File,
>>>> and allows readers to directly swap a Base File column chunk with the 
>>>> Column
>>>> File one on byte level.
>>>>
>>>> By "aligned" here I mean that if a Base File contains two row groups with 
>>>> rows
>>>> e.g. 0-1000 and 1001-2000, the Column File contains two row groups with 
>>>> exactly
>>>> the same row boundaries 0-1000 and 1001-2000; while "unaligned" allows the
>>>> Column File to contain row boundaries e.g. 0-1500 and 1501-2000.
>>>>
>>>> *Performance benefits*
>>>>
>>>> Assuming a somewhat uniform distrbution of row group byte sizes (say, RG 
>>>> is the
>>>> average row group size in bytes), reading a single Column File incurs 
>>>> anywhere
>>>> from 0 (row groups already match by chance) to 2*RG (one row before, one 
>>>> row
>>>> after) bytes of overhead I/O and memory that is being discarded by the 
>>>> reader.
>>>>
>>>> Allowing always aligning row group boundaries allows readers to lower that
>>>> overhead cost to exactly zero.
>>>>
>>>> *Pros:*
>>>>
>>>>     - Writers are not required to implement -- keeping the flag false is 
>>>> still
>>>>       valid for any data shape.
>>>>     - Readers are not required to implement -- aligning row groups doesn't 
>>>> break
>>>>       readers that still want to do stitching on row level.
>>>>     - One less copy of data along the way from parquet to executor.
>>>>     - Less reader I/O.
>>>>
>>>> *Cons:*
>>>>
>>>>     - Implementations are optional -- leads to both diverging features in
>>>>       different implementations on one hand, and diverging codepaths 
>>>> within a
>>>>       single Iceberg implementation as support for the unaligned case is 
>>>> still
>>>>       required.
>>>>     - One more field in the Column File struct -- more support work.
>>>>
>>>> Thanks!
>>>> Leonid.

Reply via email to