Hey all,

I largely agree with Anoop here, the format largely has biased towards
cheap upgrades, where readers just have to be able to handle the older
structures as well. The spec defines how to assign first row IDs for older
manifest entries etc. I do think that this is the right tradeoff to
continue making but we can always have implementations which do "eager"
upgrades on write, which I think is a good idea for "not huge" tables.

In other words, while the format *enables* cheap upgrades, it's not a
requirement for implementations to *have to perform* a "cheap" upgrade.

I do think it's valid to jump from v1/v2 to v4; I can't see an obvious
reason to prevent that.

A few syncs ago, we also talked about having leaf manifests for v3 and
older manifests in a v4 root manifest, having the ability to also reference
manifest DVs. This was so that for large updates/deletes which touch older
manifests, there's not a significant post-upgrade performance cliff from
metadata write amplification when trying to produce the combined entry
structure that v4 requires. This is more of an issue when many delete
manifests are touched, because finding the corresponding data file entries
is a potentially large scan over larger data manifests.

I think the concerns around having a proliferation of older manifests are
somewhat valid but practically I think with nearly every write pattern
there will be ways to age out the older manifest structures, either on the
commit path of an implementation or just in the background as part of
rewrite manifests. In the update/delete case I referenced, at some point
we'd need to compact the MDV's anyways for those older manifests, and an
implementation could produce the combined entries for those manifests at
that point where the cost of producing combined entry is reduced. In an
append only case, older manifests would largely just need to be rewritten
into newer manifests (probably in the background), and there's no 2 phase
planning that's needed anyways in this case.

Again though, all of the above are just implementation choices and I think
we have all the tools to address these cases; my recommendation is that we
keep the spec open enough to enable implementations to move into the newer
structures, whenever and however they choose. I don't think the spec needs
to define anything here so long as we know post-upgrade there's no
degenerate cases that we don't have a path through.

Thanks,

Amogh Jahagirdar

On Tue, Aug 11, 2026 at 11:03 PM <[email protected]> wrote:

> Hi Shawn & Anoop,
>
> Thank you Shawn for the clear write up. I am also worried that if we don’t
> provide new ways to more intentionally upgrade tables sufficiently
> significantly large pre-v4 tables may never fully complete their migrations.
>
> This concern for me personally is mostly forward looking. I am a little
> anxious about what corners that would put us into come Table format v6 or
> v7. The scenario I’m imaging is where query performance is no longer
> consistent because a user will never know if their query is about to hit a
> v3, or v4, or v5, or v6, etc etc manifest/data file.
>
> With Iceberg definitely becoming a solution for multi Petabyte sized
> tables and ideally aiming to support multi Exabyte scale tables. Waiting
> for the current maintenance options to rewrite (compact) files leaves us
> with few options for more intentional, incremental and asynchronous ways of
> catching a table up to the intended format version for all files.
>
> Anoop, question because I missed a few of the AMT syncs. Have we discussed
> if it will be valid to jump straight from V1/V2 to V4 in the proposed
> reference implementation or will users need to go V3 to V4? I’m expecting
> that it would be valid for a V1 -> V4 jump, but I don’t want to assume.
>
> Best Regards,
> Kurtis
>
> On Aug 11, 2026, at 18:51, Anoop Johnson <[email protected]> wrote:
>
> 
> Hi, Shawn -
>
> Thanks for starting the discussion on this. Your summary of our current
> thinking on upgrades is accurate. The upgrade is an O(1) operation that
> leaves prior manifest files intact. New table writes will produce the v4
> metadata. This is by design because table upgrades that require expensive
> metadata rewrites add a lot of friction to upgrade to v4. Indeed, tables
> with older manifests don't get the performance benefits of v4, but they
> produce the correct results and are no worse in performance than the v3
> table pre-upgrade.  The table converges to the v4 metadata representation
> over time as older data age out and maintenance kicks in.
>
> This is no different than any existing version upgrade: v1 to v2 or v2 to
> v3 upgrades didn't require expensive metadata or data rewrites. V4 doesn't
> make that upgrade process any worse.
>
> On the spec vs guidance. I lean towards guidance plus maintenance
> utilities rather than a prescriptive upgrade protocol that forces
> metadata rewrites.
>
> Best,
> Anoop
>
> On Tue, Aug 11, 2026 at 5:45 PM Shawn Chang <[email protected]>
> wrote:
>
>> Hi all,
>>
>> As V4 takes better shape daily, I think some items regarding the upgrade
>> path remain unclear.
>>
>> I'd like to better understand the expected V3 -> V4 upgrade path,
>> particularly the operational expectations around pre-V4 manifests after a
>> table is upgraded.
>>
>> From the current V4 work, my understanding is that an implementation can
>> perform a lightweight upgrade by creating a V4 root manifest that
>> references existing pre v4 manifests, and then writes new metadata in the
>> V4 format going forward.
>>
>> What I'm less clear about is whether we have an expected or recommended
>> lifecycle for those pre-V4 manifests after the upgrade.
>>
>> For example, is the expectation that an upgraded table may indefinitely
>> retain the structure below, until the user explicitly runs a manifest
>> rewrite/maintenance job?
>>
>> V4 root
>>   -> pre-V4 data/delete manifests
>>   -> V4 manifests
>>
>> If so, I worry that we may be shifting significant migration
>> responsibility onto users. In practice, users often don't run optional
>> table maintenance consistently. Leaving this undefined reminds me of
>> equality deletes.
>>
>> In short, I believe we can either 1) spec it or 2) provide general
>> guidance to avoid a situation where the format technically supports a clean
>> migration path, but the practical path relies on users understanding V4.
>>
>> Have we considered solutions where the implementation eagerly converts
>> metadata that is cheap to convert while explicitly leaving legacy data that
>> is expensive to migrate alone?(applying EqDels)
>>
>>
>> Best,
>>
>> Shawn
>>
>>

Reply via email to