Ah, I missed `splitOffsets`, thanks! That really does solve both the planning and the execution problem.
Would this be worth mentioning in the design doc? On Thu, Aug 13, 2026 at 7:22 AM Péter Váry <[email protected]> wrote: > > As I wrote earlier, I also see the flag as unnecessary, but having the > possibility of writing rowgroup aligned files would be nice to have in the > reference implementation. That said, I would prefer to keep this optional, as > in some cases it would be beneficial not to align the files. > > On Wed, Aug 12, 2026, 22:27 Russell Spitzer <[email protected]> wrote: >> >> I think Anoop addressed that. We already (optionally) store split offsets, >> which tell scan planning how to break a file into tasks so different tasks >> can read different parts of the same file. For Parquet, those are the >> row-group offsets in the file, so a task is effectively “read bytes X–Y” and >> can cover a single row group. >> >> At the reader, that is not a different codepath from reading the whole file. >> The reader still opens the footer, maps the task’s byte range to row >> group(s), and only reads those groups. With a column file, the extra step >> is: from the base file row group, take the row indices, open the column-file >> footer, and select whichever CF row groups cover those rows. >> >> So alignment is still decided after both footers are available. A metadata >> flag does not change how we plan splits or which CF ranges a task needs. >> >> On Wed, Aug 12, 2026 at 3:08 PM Leonid Lygin via dev >> <[email protected]> wrote: >>> >>> Thanks for the comments! >>> >>> Don't we also want to know whether the row groups align during scan >>> planning? If they align, we can assume that a task would read >>> significantly less bytes than if they don't. >>> >>> On Wed, Aug 12, 2026 at 7:58 PM Anoop Johnson <[email protected]> wrote: >>> > >>> > I don't see any obvious advantage of keeping track of row group alignment >>> > in the column file metadata. As Russell pointed out, the readers must >>> > read the Parquet footers anyway before any decoding can happen. At that >>> > point, you can derive whether the row groups are aligned by simple >>> > integer comparison. Once the reader figures out the alignment, then the >>> > optimization Leonid mentioned can still work? >>> > >>> > The closest precedent is where we keep track of the `split_offsets` in >>> > the metadata. But that is warranted because we would like to know the >>> > offsets at scan planning time before any Parquet footers are actually >>> > opened. That too is an optional optimization - if the split_offsets are >>> > absent, we fall back to fixed block split computation. >>> > >>> > Best, >>> > Anoop >>> > >>> > On Wed, Aug 12, 2026 at 9:06 AM Russell Spitzer >>> > <[email protected]> wrote: >>> >> >>> >> Before adding format-specific performance optimizations like this, I >>> >> think we should ensure we are solving a real problem and that it is only >>> >> solvable at the metadata layer. >>> >> >>> >> In this case, can’t the reader tell at read time whether row groups are >>> >> aligned by comparing footers? The reader must be given both the Base >>> >> File (BF) and Column File (CF) regardless of alignment, and must open >>> >> both footers before decoding. At that point it already knows whether BF >>> >> and CF share the same row-group row boundaries, and can plan further I/O >>> >> accordingly (1:1 chunk reads when aligned; overlapping CF row groups >>> >> when not). >>> >> >>> >> Do we get a material benefit from knowing alignment ahead of time in >>> >> Iceberg metadata, versus deriving it from the footers we have to read >>> >> anyway? >>> >> >>> >> On Wed, Aug 12, 2026 at 10:44 AM Péter Váry >>> >> <[email protected]> wrote: >>> >>> >>> >>> Hi Leonid, >>> >>> >>> >>> Thanks for the proposal and for continuing the discussion. >>> >>> >>> >>> I have a couple of questions: >>> >>> >>> >>> Do you have any suggestions for how writers could reliably produce >>> >>> aligned row groups? In particular, how would an update writer obtain >>> >>> the base file's row-group boundaries in practice? Would all writers be >>> >>> expected to read the base file footer? Also, is there an easy way to >>> >>> implement this using the current Java Parquet writer APIs? This sounds >>> >>> like a nice feature. >>> >>> Do readers actually need to know in advance that row groups are >>> >>> aligned? It seems a reader could use essentially the same algorithm for >>> >>> both aligned and unaligned files: seek to the nearest row group, >>> >>> discard any rows prior to the desired starting position, and continue >>> >>> reading from there. If page-level skipping is available, it may even >>> >>> avoid reading some of those unnecessary pages. When the row group >>> >>> already begins at the required offset, the discard step simply becomes >>> >>> a no-op. If that is the case, do we actually need an alignment flag at >>> >>> all? >>> >>> >>> >>> Thanks, Peter >>> >>> >>> >>> Leonid Lygin via dev <[email protected]> ezt írta (időpont: 2026. >>> >>> aug. 12., Sze, 17:06): >>> >>>> >>> >>>> Hi all! >>> >>>> >>> >>>> Following up on the Column File representation thread >>> >>>> <https://lists.apache.org/thread/jbh1gbrso5h6l4by9rh9poy2cjjtb8j0>, >>> >>>> I'd like to >>> >>>> fire off a discussion about a possible optimization for Column >>> >>>> Updates, where >>> >>>> supporting writers might decide to write out the Column File aligning >>> >>>> all row >>> >>>> groups with the Base File, enabling supporting readers to do simple >>> >>>> zero-copy >>> >>>> reads. >>> >>>> >>> >>>> *Context* >>> >>>> >>> >>>> The original thread settled on a dense representation (i.e. Column >>> >>>> Files contain >>> >>>> exactly the same amount of rows as Base Files), while allowing >>> >>>> unsynchronized >>> >>>> row-group boundaries. >>> >>>> >>> >>>> I'm proposing adding a separate metadata flag (e.g. >>> >>>> `ColumnFile.containsAlignedRowGroups`) which, when set, indicates that >>> >>>> the row >>> >>>> groups contained in the associated Column File are aligned with the >>> >>>> Base File, >>> >>>> and allows readers to directly swap a Base File column chunk with the >>> >>>> Column >>> >>>> File one on byte level. >>> >>>> >>> >>>> By "aligned" here I mean that if a Base File contains two row groups >>> >>>> with rows >>> >>>> e.g. 0-1000 and 1001-2000, the Column File contains two row groups >>> >>>> with exactly >>> >>>> the same row boundaries 0-1000 and 1001-2000; while "unaligned" allows >>> >>>> the >>> >>>> Column File to contain row boundaries e.g. 0-1500 and 1501-2000. >>> >>>> >>> >>>> *Performance benefits* >>> >>>> >>> >>>> Assuming a somewhat uniform distrbution of row group byte sizes (say, >>> >>>> RG is the >>> >>>> average row group size in bytes), reading a single Column File incurs >>> >>>> anywhere >>> >>>> from 0 (row groups already match by chance) to 2*RG (one row before, >>> >>>> one row >>> >>>> after) bytes of overhead I/O and memory that is being discarded by the >>> >>>> reader. >>> >>>> >>> >>>> Allowing always aligning row group boundaries allows readers to lower >>> >>>> that >>> >>>> overhead cost to exactly zero. >>> >>>> >>> >>>> *Pros:* >>> >>>> >>> >>>> - Writers are not required to implement -- keeping the flag false >>> >>>> is still >>> >>>> valid for any data shape. >>> >>>> - Readers are not required to implement -- aligning row groups >>> >>>> doesn't break >>> >>>> readers that still want to do stitching on row level. >>> >>>> - One less copy of data along the way from parquet to executor. >>> >>>> - Less reader I/O. >>> >>>> >>> >>>> *Cons:* >>> >>>> >>> >>>> - Implementations are optional -- leads to both diverging features >>> >>>> in >>> >>>> different implementations on one hand, and diverging codepaths >>> >>>> within a >>> >>>> single Iceberg implementation as support for the unaligned case >>> >>>> is still >>> >>>> required. >>> >>>> - One more field in the Column File struct -- more support work. >>> >>>> >>> >>>> Thanks! >>> >>>> Leonid.
