Hi all, We discussed this in the Efficient Column Updates sync today and concluded that the _pos field will not be required, nor will the spec suggest that writers may materialize it. In other words, writing the _pos field will not be allowed.
~ Anurag On Thu, Sep 3, 2026 at 12:31 PM Marco Kroll <[email protected]> wrote: > I was thinking more about this today and I agree with Russell, this is an > engine/writer responsibility and does not need to be codified in the spec. > > On Wed, Sep 2, 2026 at 9:40 PM Russell Spitzer <[email protected]> > wrote: > >> I still favor removing it. Wouldn't any writer be able to debug this at >> write time? If the origin row position doesn't match the actual position in >> the file the writer should just break at that point, no need to persist it >> imho. >> >> On Wed, Sep 2, 2026 at 2:30 PM Marco Kroll <[email protected]> >> wrote: >> >>> Thanks everyone for chiming in. >>> I'm not too well versed in the history of Iceberg. So I searched the >>> list and the spec for precedent on storing information purely for debugging >>> purposes. >>> As far as I can tellthere is none. The closest I found is the rule that >>> all columns must be written to data files even when redundant with manifest >>> metadata, which is explicitly a backup "in case of corruption or bugs in >>> the metadata layer". >>> Please correct me if I'm wrong here. If not, this means requiring the >>> _pos column for debugging purposes only would be a first. >>> >>> I lean towards removing it. But if that is not possible, I’d prefer to >>> build on top of an earlier proposal from Andrei and would suggest making it >>> optional. >>> Then we could state following requirements: >>> - column files MUST adhere to the same positional alignment as the base >>> file >>> - _pos is never needed to resolve a read >>> - writers MAY materialize it, for example when they cannot guarantee >>> alignment, or for troubleshooting >>> - when absent, readers derive the position from the row ordinal >>> - when present and disagreeing with the ordinal, the file is corrupt, >>> rather than there being a second legal read path >>> >>> Best >>> Marco >>> >>> >>> On Fri, Aug 28, 2026 at 10:00 PM Ryan Blue <[email protected]> wrote: >>> >>>> Conceptually the value we need is not the row's position in the column >>>> update file, it is the row's position in the original data file. We expect >>>> them to be equal, but it's an unrecoverable and undetectable problem if you >>>> don't store the row position and a writer omits deleted rows. I think >>>> that's the debugging that Marco mentioned. >>>> >>>> I'd prefer to keep this since I think that Parquet should be able to >>>> store these efficiently. But that's not what happens today so this isn't a >>>> strongly held opinion. >>>> >>>> Ryan >>>> >>>> On Thu, Aug 27, 2026 at 6:54 AM Leonid Lygin via dev < >>>> [email protected]> wrote: >>>> >>>>> Glad to see that everyone (up to this point) agrees! >>>>> >>>>> Should we ratify the decision in the doc itself? >>>>> >>>>> On Wed, Aug 26, 2026 at 6:15 PM Anurag Mantripragada >>>>> <[email protected]> wrote: >>>>> > >>>>> > The addition of `_pos` column was not a remnant of previous design, >>>>> but a conscious choice we made during one of the column updates sync. I'm >>>>> in favor of removing it if it makes implementation easier. >>>>> > >>>>> > ~ Anurag >>>>> > >>>>> > On Wed, Aug 26, 2026 at 7:31 AM Gábor Kaszab <[email protected]> >>>>> wrote: >>>>> >> >>>>> >> Yes, the question is whether we want to persist _pos into the >>>>> column files. We earlier concluded that even though we went with the dense >>>>> representation we persist _pos. However, I think re-opening the question >>>>> is >>>>> reasonable, because that's just extra noise ATM, and we shouldn't write >>>>> that field into column files. Reading the _pos column would still work >>>>> regardless if we include the base file or not to the query. >>>>> >> >>>>> >> Best Regards, >>>>> >> Gabor >>>>> >> >>>>> >> >>>>> >> Russell Spitzer <[email protected]> ezt írta (időpont: >>>>> 2026. aug. 26., Sze, 16:21): >>>>> >>> >>>>> >>> What is the actual argument here? I think having a persisted field >>>>> doesn't make sense since we expect perfect alignment. We would expect >>>>> reading the file in isolation with the metadata _pos column should still >>>>> work right? >>>>> >>> >>>>> >>> If we are just discussing removing a persisted value, I'm in favor >>>>> of that. >>>>> >>> >>>>> >>> On Wed, Aug 26, 2026 at 8:37 AM Gábor Kaszab < >>>>> [email protected]> wrote: >>>>> >>>> >>>>> >>>> I hear you, and I share the same opinion. If we don't need such a >>>>> field then it's just extra unnecessary complexity to write it. I'm not >>>>> entirely convinced on the debugging use of the _pos field. Would be >>>>> beneficial to reduce unnecessary noise and confusion by not adding the >>>>> _pos >>>>> field. >>>>> >>>> >>>>> >>>> Let's discuss this on the next sync! In the meantime, opinions >>>>> are welcome here too. >>>>> >>>> >>>>> >>>> Thanks, >>>>> >>>> Gabor >>>>> >>>> >>>>> >>>> Andrei Tserakhau via dev <[email protected]> ezt írta >>>>> (időpont: 2026. aug. 26., Sze, 15:32): >>>>> >>>>> >>>>> >>>>> +1 on this question. >>>>> >>>>> >>>>> >>>>> Right now `_pos` column feels more like debug leftovers, it >>>>> bring some confusion for read-side weather it's expected to be readed or >>>>> not. >>>>> >>>>> >>>>> >>>>> I think removing it would make implementation easier. >>>>> >>>>> >>>>> >>>>> Best, >>>>> >>>>> Andrei >>>>> >>>>> >>>>> >>>>> On Wed, Aug 26, 2026 at 2:42 PM Leonid Lygin via dev < >>>>> [email protected]> wrote: >>>>> >>>>>> >>>>> >>>>>> Thanks for the quick response! >>>>> >>>>>> >>>>> >>>>>> My biggest concern with `_pos` is not performance but rather >>>>> clarity >>>>> >>>>>> and implementation divergence: >>>>> >>>>>> >>>>> >>>>>> 1. including `_pos` is redundant, and (at least for me) >>>>> provokes a >>>>> >>>>>> re-read of the row alignment section — "why include `_pos` if >>>>> files >>>>> >>>>>> are fully aligned?"; >>>>> >>>>>> 2. having `_pos` fully duplicate the row position, there are two >>>>> >>>>>> different legal ways to implement reads: either positionally, >>>>> or using >>>>> >>>>>> `_pos`. >>>>> >>>>>> >>>>> >>>>>> On Wed, Aug 26, 2026 at 2:35 PM Gábor Kaszab < >>>>> [email protected]> wrote: >>>>> >>>>>> > >>>>> >>>>>> > Hey All, >>>>> >>>>>> > >>>>> >>>>>> > Thanks for bringing this up! (for me the initial mail went to >>>>> spam, though...) >>>>> >>>>>> > >>>>> >>>>>> > Technically, with the dense representation we don't really >>>>> need the _pos column in the column files, unless for troubleshooting. >>>>> While >>>>> checking the row counts is good, if they don't match we might get a better >>>>> understanding on what the writer missed writing if we had the _pos col, >>>>> also the order could be verified. >>>>> >>>>>> > >>>>> >>>>>> > Apart from debugging, I think either way is just fine. An >>>>> additional detail to consider is that according to my experiments, there >>>>> isn't really any storage cost for writing the _pos with delta encoding >>>>> (e.g. with Parquet V2). So the conclusion was that since it comes for >>>>> free, >>>>> and might help for debugging, why not write it. >>>>> >>>>>> > >>>>> >>>>>> > Should we reopen this question? Any further feedback is >>>>> welcome. >>>>> >>>>>> > >>>>> >>>>>> > Best Regards, >>>>> >>>>>> > Gabor >>>>> >>>>>> > >>>>> >>>>>> > Leonid Lygin via dev <[email protected]> ezt írta >>>>> (időpont: 2026. aug. 26., Sze, 14:10): >>>>> >>>>>> >> >>>>> >>>>>> >> Definitely agree that including `_pos` raises questions. >>>>> >>>>>> >> >>>>> >>>>>> >> If "debugging" is to be understood as figuring out if the >>>>> column files >>>>> >>>>>> >> have gaps -- just checking the row counts is good enough for >>>>> that. Is >>>>> >>>>>> >> there a lot to be gained from figuring out where exactly the >>>>> gap is >>>>> >>>>>> >> occurring? >>>>> >>>>>> >> >>>>> >>>>>> >> On Mon, Aug 24, 2026 at 1:57 PM Marco Kroll >>>>> >>>>>> >> <[email protected]> wrote: >>>>> >>>>>> >> > >>>>> >>>>>> >> > Hi all, >>>>> >>>>>> >> > >>>>> >>>>>> >> > I just saw the agenda [1] for tomorrow's (2026-08-25) sync >>>>> and want to +1 the `_pos` column topic. >>>>> >>>>>> >> > My understanding is that this column exists for two >>>>> reasons: >>>>> >>>>>> >> > 1. debugging >>>>> >>>>>> >> > 2. detect if writers skipped deleted rows >>>>> >>>>>> >> > >>>>> >>>>>> >> > My take is that using the dense Null filled representation >>>>> addresses both of these issues. >>>>> >>>>>> >> > It implicitly encodes the position, very much like for >>>>> deletion vectors and since all rows need to be present, comparing the row >>>>> count of the base file with the column file can be used to verify that all >>>>> rows were written. >>>>> >>>>>> >> > >>>>> >>>>>> >> > The main thing to add to the doc would be that the row >>>>> order must be identical to the base file. >>>>> >>>>>> >> > >>>>> >>>>>> >> > Best >>>>> >>>>>> >> > Marco >>>>> >>>>>> >> > >>>>> >>>>>> >> > [1]: >>>>> https://docs.google.com/document/d/1Bd7JVzgajA8-DozzeEE24mID_GLuz6iwj0g4TlcVJcs/edit?tab=t.jvm7iiiulf8q#heading=h.rbisiun18esp >>>>> >>>>
