I think Anoop addressed that. We already (optionally) store split offsets, which tell scan planning how to break a file into tasks so different tasks can read different parts of the same file. For Parquet, those are the row-group offsets in the file, so a task is effectively “read bytes X–Y” and can cover a single row group.
At the reader, that is not a different codepath from reading the whole file. The reader still opens the footer, maps the task’s byte range to row group(s), and only reads those groups. With a column file, the extra step is: from the base file row group, take the row indices, open the column-file footer, and select whichever CF row groups cover those rows. So alignment is still decided after both footers are available. A metadata flag does not change how we plan splits or which CF ranges a task needs. On Wed, Aug 12, 2026 at 3:08 PM Leonid Lygin via dev <[email protected]> wrote: > Thanks for the comments! > > Don't we also want to know whether the row groups align during scan > planning? If they align, we can assume that a task would read > significantly less bytes than if they don't. > > On Wed, Aug 12, 2026 at 7:58 PM Anoop Johnson <[email protected]> wrote: > > > > I don't see any obvious advantage of keeping track of row group > alignment in the column file metadata. As Russell pointed out, the readers > must read the Parquet footers anyway before any decoding can happen. At > that point, you can derive whether the row groups are aligned by simple > integer comparison. Once the reader figures out the alignment, then the > optimization Leonid mentioned can still work? > > > > The closest precedent is where we keep track of the `split_offsets` in > the metadata. But that is warranted because we would like to know the > offsets at scan planning time before any Parquet footers are actually > opened. That too is an optional optimization - if the split_offsets are > absent, we fall back to fixed block split computation. > > > > Best, > > Anoop > > > > On Wed, Aug 12, 2026 at 9:06 AM Russell Spitzer < > [email protected]> wrote: > >> > >> Before adding format-specific performance optimizations like this, I > think we should ensure we are solving a real problem and that it is only > solvable at the metadata layer. > >> > >> In this case, can’t the reader tell at read time whether row groups are > aligned by comparing footers? The reader must be given both the Base File > (BF) and Column File (CF) regardless of alignment, and must open both > footers before decoding. At that point it already knows whether BF and CF > share the same row-group row boundaries, and can plan further I/O > accordingly (1:1 chunk reads when aligned; overlapping CF row groups when > not). > >> > >> Do we get a material benefit from knowing alignment ahead of time in > Iceberg metadata, versus deriving it from the footers we have to read > anyway? > >> > >> On Wed, Aug 12, 2026 at 10:44 AM Péter Váry < > [email protected]> wrote: > >>> > >>> Hi Leonid, > >>> > >>> Thanks for the proposal and for continuing the discussion. > >>> > >>> I have a couple of questions: > >>> > >>> Do you have any suggestions for how writers could reliably produce > aligned row groups? In particular, how would an update writer obtain the > base file's row-group boundaries in practice? Would all writers be expected > to read the base file footer? Also, is there an easy way to implement this > using the current Java Parquet writer APIs? This sounds like a nice feature. > >>> Do readers actually need to know in advance that row groups are > aligned? It seems a reader could use essentially the same algorithm for > both aligned and unaligned files: seek to the nearest row group, discard > any rows prior to the desired starting position, and continue reading from > there. If page-level skipping is available, it may even avoid reading some > of those unnecessary pages. When the row group already begins at the > required offset, the discard step simply becomes a no-op. If that is the > case, do we actually need an alignment flag at all? > >>> > >>> Thanks, Peter > >>> > >>> Leonid Lygin via dev <[email protected]> ezt írta (időpont: > 2026. aug. 12., Sze, 17:06): > >>>> > >>>> Hi all! > >>>> > >>>> Following up on the Column File representation thread > >>>> <https://lists.apache.org/thread/jbh1gbrso5h6l4by9rh9poy2cjjtb8j0>, > I'd like to > >>>> fire off a discussion about a possible optimization for Column > Updates, where > >>>> supporting writers might decide to write out the Column File aligning > all row > >>>> groups with the Base File, enabling supporting readers to do simple > zero-copy > >>>> reads. > >>>> > >>>> *Context* > >>>> > >>>> The original thread settled on a dense representation (i.e. Column > Files contain > >>>> exactly the same amount of rows as Base Files), while allowing > unsynchronized > >>>> row-group boundaries. > >>>> > >>>> I'm proposing adding a separate metadata flag (e.g. > >>>> `ColumnFile.containsAlignedRowGroups`) which, when set, indicates > that the row > >>>> groups contained in the associated Column File are aligned with the > Base File, > >>>> and allows readers to directly swap a Base File column chunk with the > Column > >>>> File one on byte level. > >>>> > >>>> By "aligned" here I mean that if a Base File contains two row groups > with rows > >>>> e.g. 0-1000 and 1001-2000, the Column File contains two row groups > with exactly > >>>> the same row boundaries 0-1000 and 1001-2000; while "unaligned" > allows the > >>>> Column File to contain row boundaries e.g. 0-1500 and 1501-2000. > >>>> > >>>> *Performance benefits* > >>>> > >>>> Assuming a somewhat uniform distrbution of row group byte sizes (say, > RG is the > >>>> average row group size in bytes), reading a single Column File incurs > anywhere > >>>> from 0 (row groups already match by chance) to 2*RG (one row before, > one row > >>>> after) bytes of overhead I/O and memory that is being discarded by > the reader. > >>>> > >>>> Allowing always aligning row group boundaries allows readers to lower > that > >>>> overhead cost to exactly zero. > >>>> > >>>> *Pros:* > >>>> > >>>> - Writers are not required to implement -- keeping the flag false > is still > >>>> valid for any data shape. > >>>> - Readers are not required to implement -- aligning row groups > doesn't break > >>>> readers that still want to do stitching on row level. > >>>> - One less copy of data along the way from parquet to executor. > >>>> - Less reader I/O. > >>>> > >>>> *Cons:* > >>>> > >>>> - Implementations are optional -- leads to both diverging > features in > >>>> different implementations on one hand, and diverging codepaths > within a > >>>> single Iceberg implementation as support for the unaligned case > is still > >>>> required. > >>>> - One more field in the Column File struct -- more support work. > >>>> > >>>> Thanks! > >>>> Leonid. >
