Steven, is there an issue or PR related to the monotonic snapshot timestamp and a new *_last_updated_timestamp_ms* metadata column and/or doc that we can link to an improvement issue? I know there's been discussion, but I don't know where that currently stands.
Manu, apologies for the confusion regarding Efficient Column Updates and Column Appends, I consider them the same (feature names tend to evolve over time like SFC/AMT), but the former is a more accurate name. Kurtis, I think there's a lot of interest in correctness testing, integrity checks, and cross validation, but I'm not sure that's necessarily a spec blocker (probably a good point of discussion as we move closer to a v4). A lot of people brought up the collations discussion, which is great. I've asked Andrei to create a spec improvement issue for that as well. I have started to update the existing V4 Milestone <https://github.com/apache/iceberg/milestone/58> and will update it as we follow up with proposal issues (per the proposal process <https://iceberg.apache.org/contribute/#apache-iceberg-improvement-proposals> ). I would also note that many supporting changes are likely necessary, so I don't expect this to be an exhaustive detailed list at this point, but please continue to highlight anything we may have missed. -Dan On Wed, Aug 12, 2026 at 4:21 AM Andrei Tserakhau via dev < [email protected]> wrote: > Thanks for kicking this off, Dan. Two items I’d add to the v4 scope. > > == Collation == > > For v4 scoping: should we include collation? > > Steven and Kurtis both flagged it. There’s an active thread off Loeser’s > proposal, > spec PR apache/iceberg#16972, and an Aug 5 design sync with Steven, > Russell, > and Alexander that aligned on the feature boundaries. > > It fits v4 because it depends on v4-only machinery: per-collation bounds as > generated-expression column stats and a file-level ICU-version tag. It also > intersects with the equality-delete deprecation Manu raised: as Russell > pointed out, equality deletes would not apply to collated columns. > > The sync also settled the main compatibility points: collation is optional, > readers can fall back to non-collated stats, ICU compatibility is > engine-owned, and collated columns cannot be used for equality deletes or > partitioning. Bounds use the original strings. > > Alexander is folding this into the proposal update; details are in #16972. > > I’m happy to continue to drive the tracking of it. > > == Efficient column updates / Column Append == > > On Manu’s question: I read Column Append and efficient column updates as > two operations over the same column-file representation. One appends values > for a new column; the other replaces values for an existing one. In both > cases, the values live in a separate file aligned with the base data file > rather than requiring a full copy-on-write rewrite. > > If that’s the intent, I’d track them together around the shared > representation rather than as separate v4 items. > > There are already two active threads here: Anurag’s efficient column > updates discussion and the Column Update File Representation thread, which > is getting close to a decision on positional alignment. > > One cross-format consideration: we’ve been working on the same problem > from the Delta side, and the write-path prototype converged on the same > dense, row-aligned representation. It would be useful to keep the > representation compatible across the two formats while it is still being > defined. > > == Tracking == > > +1 on GitHub milestones + projects, same as the last release. > > Best, > Andrei > > On Wed, Aug 12, 2026 at 7:38 AM <[email protected]> wrote: > >> Hi Daniel, >> >> Thank you for starting the thread and laying out current work in >> consideration for V4. I agree with the current list and would like to add >> other ideas below. >> >> I think there is still space for formalizing a specification language on >> how to implement and integrate correctness and integrity checks for Iceberg. >> >> I am curious on the discussion around collation that Steven called out. >> >> Best Regards, >> Kurtis >> >> On Aug 11, 2026, at 19:51, Manu Zhang <[email protected]> wrote: >> >> >> Hi Dan, >> >> Thanks for the summary. When you say "Column Append," are you referring >> to efficient column updates[1]? >> Besides new features, Huaxin has initiated a discussion about deprecating >> equality deletes which remains unresolved[2]. >> BTW, it looks the video of the last Community Sync hasn't been >> uploaded[3]. Does anyone know about it? >> >> 1. https://www.youtube.com/watch?v=dDZKzz3PKIk >> 2. https://lists.apache.org/thread/ks01jpv40qjlvz4yop5tlqv4x5oxbwy6 >> 3. https://www.youtube.com/@ApacheIceberg/videos >> >> Thanks, >> Manu >> >> On Wed, Aug 12, 2026 at 6:16 AM Steven Wu <[email protected]> wrote: >> >>> I would like to add one item to the "Table features" that we previously >>> discussed and agreed upon [1]. >>> - Monotonic snapshot timestamp and a new *_last_updated_timestamp_ms* >>> metadata >>> column that can inherit from the snapshot timestamp. >>> >>> I know there is also an ongoing discussion on collation. >>> >>> 1. https://lists.apache.org/thread/cln7xtjy9kw2096gzqlg5c4gg8b9td10 >>> >>> >>> On Tue, Aug 11, 2026 at 2:42 PM Daniel Weeks <[email protected]> wrote: >>> >>>> Hey Everyone, >>>> >>>> At the last Iceberg Community Sync, I raised the topic of opening a >>>> discussion on how we want to scope v4 to arrive at a shared understanding >>>> of what we want to land in the next table spec version. Anyone who >>>> regularly attends conferences, meetups, or syncs is likely familiar with >>>> specific efforts currently in flight, but as a community, we haven't >>>> defined what we believe constitutes a cohesive set of capabilities to >>>> target for the next release. >>>> >>>> It's early enough at this point to start a discussion and still have >>>> time to include critical features not yet identified or push back anything >>>> that may either be underdeveloped or poorly understood. >>>> >>>> Broadly speaking, the various ongoing workstreams fall into three main >>>> categories: >>>> >>>> *Content Metadata Updates* >>>> - Adaptive Metadata Tree (AMT) / Single File Commits >>>> - Column Statistics >>>> - Relative Paths >>>> - Column Append >>>> >>>> *Table Features* >>>> - Check Constraints >>>> - Default Value Expressions >>>> - Generated Columns >>>> >>>> *Data Types* >>>> - File Type >>>> - Vector Type >>>> >>>> Some have only recently been proposed (e.g. Vector Type and Generated >>>> Columns), while others have been discussed for over a year and are well >>>> into implementation (e.g. AMT and Stats). >>>> >>>> While many non-table format discussions and efforts are ongoing (MVs, >>>> Indexing, REST), focusing attention on what aligns with the v4 release and >>>> communicating the scope and progress of that effort would be helpful. For >>>> the last release, we used a combination of GitHub milestones and projects >>>> (some exist for v4 efforts but are currently incomplete/unmaintained), >>>> which I believe many found helpful. >>>> >>>> I'd like to hear everyone's thoughts: >>>> - Does this algin with everyone's interests in v4? >>>> - Are there features we've discussed that should/shouldn't be included? >>>> - Do we want to start tracking progress across these efforts? >>>> >>>> Looking forward to discussion and feedback, >>>> -Dan >>>> >>>
