Ah, I missed `splitOffsets`, thanks! That really does solve both the
planning and the execution problem.

Would this be worth mentioning in the design doc?

On Thu, Aug 13, 2026 at 7:22 AM Péter Váry <[email protected]> wrote:
>
> As I wrote earlier, I also see the flag as unnecessary, but having the 
> possibility of writing rowgroup aligned files would be nice to have in the 
> reference implementation. That said, I would prefer to keep this optional, as 
> in some cases it would be beneficial not to align the files.
>
> On Wed, Aug 12, 2026, 22:27 Russell Spitzer <[email protected]> wrote:
>>
>> I think Anoop addressed that. We already (optionally) store split offsets, 
>> which tell scan planning how to break a file into tasks so different tasks 
>> can read different parts of the same file. For Parquet, those are the 
>> row-group offsets in the file, so a task is effectively “read bytes X–Y” and 
>> can cover a single row group.
>>
>> At the reader, that is not a different codepath from reading the whole file. 
>> The reader still opens the footer, maps the task’s byte range to row 
>> group(s), and only reads those groups. With a column file, the extra step 
>> is: from the base file row group, take the row indices, open the column-file 
>> footer, and select whichever CF row groups cover those rows.
>>
>> So alignment is still decided after both footers are available. A metadata 
>> flag does not change how we plan splits or which CF ranges a task needs.
>>
>> On Wed, Aug 12, 2026 at 3:08 PM Leonid Lygin via dev 
>> <[email protected]> wrote:
>>>
>>> Thanks for the comments!
>>>
>>> Don't we also want to know whether the row groups align during scan
>>> planning? If they align, we can assume that a task would read
>>> significantly less bytes than if they don't.
>>>
>>> On Wed, Aug 12, 2026 at 7:58 PM Anoop Johnson <[email protected]> wrote:
>>> >
>>> > I don't see any obvious advantage of keeping track of row group alignment 
>>> > in the column file metadata. As Russell pointed out, the readers must 
>>> > read the Parquet footers anyway before any decoding can happen. At that 
>>> > point, you can derive whether the row groups are aligned by simple 
>>> > integer comparison. Once the reader figures out the alignment, then the 
>>> > optimization Leonid mentioned can still work?
>>> >
>>> > The closest precedent is where we keep track of the `split_offsets` in 
>>> > the metadata. But that is warranted because we would like to know the 
>>> > offsets at scan planning time before any Parquet footers are actually 
>>> > opened. That too is an optional optimization - if the split_offsets are 
>>> > absent, we fall back to fixed block split computation.
>>> >
>>> > Best,
>>> > Anoop
>>> >
>>> > On Wed, Aug 12, 2026 at 9:06 AM Russell Spitzer 
>>> > <[email protected]> wrote:
>>> >>
>>> >> Before adding format-specific performance optimizations like this, I 
>>> >> think we should ensure we are solving a real problem and that it is only 
>>> >> solvable at the metadata layer.
>>> >>
>>> >> In this case, can’t the reader tell at read time whether row groups are 
>>> >> aligned by comparing footers? The reader must be given both the Base 
>>> >> File (BF) and Column File (CF) regardless of alignment, and must open 
>>> >> both footers before decoding. At that point it already knows whether BF 
>>> >> and CF share the same row-group row boundaries, and can plan further I/O 
>>> >> accordingly (1:1 chunk reads when aligned; overlapping CF row groups 
>>> >> when not).
>>> >>
>>> >> Do we get a material benefit from knowing alignment ahead of time in 
>>> >> Iceberg metadata, versus deriving it from the footers we have to read 
>>> >> anyway?
>>> >>
>>> >> On Wed, Aug 12, 2026 at 10:44 AM Péter Váry 
>>> >> <[email protected]> wrote:
>>> >>>
>>> >>> Hi Leonid,
>>> >>>
>>> >>> Thanks for the proposal and for continuing the discussion.
>>> >>>
>>> >>> I have a couple of questions:
>>> >>>
>>> >>> Do you have any suggestions for how writers could reliably produce 
>>> >>> aligned row groups? In particular, how would an update writer obtain 
>>> >>> the base file's row-group boundaries in practice? Would all writers be 
>>> >>> expected to read the base file footer? Also, is there an easy way to 
>>> >>> implement this using the current Java Parquet writer APIs? This sounds 
>>> >>> like a nice feature.
>>> >>> Do readers actually need to know in advance that row groups are 
>>> >>> aligned? It seems a reader could use essentially the same algorithm for 
>>> >>> both aligned and unaligned files: seek to the nearest row group, 
>>> >>> discard any rows prior to the desired starting position, and continue 
>>> >>> reading from there. If page-level skipping is available, it may even 
>>> >>> avoid reading some of those unnecessary pages. When the row group 
>>> >>> already begins at the required offset, the discard step simply becomes 
>>> >>> a no-op. If that is the case, do we actually need an alignment flag at 
>>> >>> all?
>>> >>>
>>> >>> Thanks, Peter
>>> >>>
>>> >>> Leonid Lygin via dev <[email protected]> ezt írta (időpont: 2026. 
>>> >>> aug. 12., Sze, 17:06):
>>> >>>>
>>> >>>> Hi all!
>>> >>>>
>>> >>>> Following up on the Column File representation thread
>>> >>>> <https://lists.apache.org/thread/jbh1gbrso5h6l4by9rh9poy2cjjtb8j0>, 
>>> >>>> I'd like to
>>> >>>> fire off a discussion about a possible optimization for Column 
>>> >>>> Updates, where
>>> >>>> supporting writers might decide to write out the Column File aligning 
>>> >>>> all row
>>> >>>> groups with the Base File, enabling supporting readers to do simple 
>>> >>>> zero-copy
>>> >>>> reads.
>>> >>>>
>>> >>>> *Context*
>>> >>>>
>>> >>>> The original thread settled on a dense representation (i.e. Column 
>>> >>>> Files contain
>>> >>>> exactly the same amount of rows as Base Files), while allowing 
>>> >>>> unsynchronized
>>> >>>> row-group boundaries.
>>> >>>>
>>> >>>> I'm proposing adding a separate metadata flag (e.g.
>>> >>>> `ColumnFile.containsAlignedRowGroups`) which, when set, indicates that 
>>> >>>> the row
>>> >>>> groups contained in the associated Column File are aligned with the 
>>> >>>> Base File,
>>> >>>> and allows readers to directly swap a Base File column chunk with the 
>>> >>>> Column
>>> >>>> File one on byte level.
>>> >>>>
>>> >>>> By "aligned" here I mean that if a Base File contains two row groups 
>>> >>>> with rows
>>> >>>> e.g. 0-1000 and 1001-2000, the Column File contains two row groups 
>>> >>>> with exactly
>>> >>>> the same row boundaries 0-1000 and 1001-2000; while "unaligned" allows 
>>> >>>> the
>>> >>>> Column File to contain row boundaries e.g. 0-1500 and 1501-2000.
>>> >>>>
>>> >>>> *Performance benefits*
>>> >>>>
>>> >>>> Assuming a somewhat uniform distrbution of row group byte sizes (say, 
>>> >>>> RG is the
>>> >>>> average row group size in bytes), reading a single Column File incurs 
>>> >>>> anywhere
>>> >>>> from 0 (row groups already match by chance) to 2*RG (one row before, 
>>> >>>> one row
>>> >>>> after) bytes of overhead I/O and memory that is being discarded by the 
>>> >>>> reader.
>>> >>>>
>>> >>>> Allowing always aligning row group boundaries allows readers to lower 
>>> >>>> that
>>> >>>> overhead cost to exactly zero.
>>> >>>>
>>> >>>> *Pros:*
>>> >>>>
>>> >>>>     - Writers are not required to implement -- keeping the flag false 
>>> >>>> is still
>>> >>>>       valid for any data shape.
>>> >>>>     - Readers are not required to implement -- aligning row groups 
>>> >>>> doesn't break
>>> >>>>       readers that still want to do stitching on row level.
>>> >>>>     - One less copy of data along the way from parquet to executor.
>>> >>>>     - Less reader I/O.
>>> >>>>
>>> >>>> *Cons:*
>>> >>>>
>>> >>>>     - Implementations are optional -- leads to both diverging features 
>>> >>>> in
>>> >>>>       different implementations on one hand, and diverging codepaths 
>>> >>>> within a
>>> >>>>       single Iceberg implementation as support for the unaligned case 
>>> >>>> is still
>>> >>>>       required.
>>> >>>>     - One more field in the Column File struct -- more support work.
>>> >>>>
>>> >>>> Thanks!
>>> >>>> Leonid.

Reply via email to