Hi all,

I incorporated the results from our meeting last week into a v2 of the design doc. The new version can be found in the same doc <https://docs.google.com/document/d/1m8b7u97uteHYjXk-4DNglJSpQO8OcZOCzW2tApCNTW4/edit?tab=t.50g4t7yl7w7t#heading=h.td5on8bji7s3> as the original proposal, in a separate "v2" tab.

On top of the points we agreed on, it also addresses some of Andrei's open comments on v1 - e.g., it has more details on supported specifiers, including a formal grammar.

The remaining major open points are:
- how to best integrate the metrics into the spec (including schema evolution/changing collations)
- which stats to store to maximize cross-version pruning capabilities

If you get the chance, please take a look. I'll be off next week but will look at feedback once I'm back from vacation.

Best,
Alex

On 8/7/26 19:26, Alexander Löser wrote:

Hi all,

we met on Wednesday, and reached an alignment on the direction this proposal should take.

Specifically, these are the directions we agreed on:

  * The ICU version should remain an engine decision; the spec will
    not mandate a specific version. Implications:
      o Writers must tag the bounds they write with the ICU version
        they used
      o Readers must only use bounds written with an ICU version they
        can handle
  * The collation bounds will be represented using original strings
    rather than collation keys
      o on top of the bounds, we will collect an additional metric
        that provides insights over the code points present in a file,
        which can enable engines to perform cross-ICU version pruning
      o the exact shape of that metric still needs to be determined -
        a very simple case could be a flag indicating the file
        contains only ASCII
          + It will be the engine's responsibility to decide whether
            the bounds can be reused across versions. The spec only
            provides information on the code points present in the file
  * Equality deletes will either be deprecated in V4, or not work on
    collated columns

There are some open points that require additional investigation:

  * How should the ICU version be encoded in the bounds? One generated
    expression per collated column per ICU version present might lead
    to a schema explosion
  * How should a code-point related metric look like to enable as much
    cross-ICU version pruning as possible?

I've started updating the original proposal to incorporate the direction we've aligned on. I will also include the tradeoffs we considered and the reasons for our decisions.
I'm planning to share the updated proposal some time next week.

Best,
Alex

On 7/17/26 22:30, Andrei Tserakhau via dev wrote:
Sure!

Scheduled at 5th of August, 5 PM CET / 8 AM PST

Dedicated sync on collation support for the Iceberg spec (PR #16972).

*Goal:* work through the open decisions so we can either commit to a v1 shape or agree it's not worth pursuing.

Decisions to make (in order):

1. *Do we pursue this at all? *Is column-level collation worth adding to the Iceberg spec, given the complexity below, or is it better left to engines? Everything else is moot if the answer is no. 2. *Who owns the ICU version: *format or engine? Pin one version in the format (deterministic across engines, unambiguous equality deletes, but lockstep upgrades) vs. let each engine own its version (independent upgrades, but the same query can return different results on different engines). May split by layer: pruning (performance) vs execution semantics (correctness). 3. *Equality deletes on collated columns:* allow, or disallow in v1? Different ICU versions can disagree whether a delete matched, which can change results for later queries even on non-collated columns. 4. *v1 scope:* pruning + annotation only, or execution semantics too? What's explicitly out (sort orders, partition/bucket transforms)? 5. *Provider model:* restricted registered set like geo (icu to start), or open namespace?

This binds every engine that implements collations, so we want implementers in the room before we fix field IDs. Please come with your engine's ICU-version story: pinned, versionless, and upgrade cadence.

Pre-read:
- Spec PR: iceberg#16972 <https://github.com/apache/iceberg/pull/16972>
- Original proposal: Iceberg Proposal - Collation Support <https://docs.google.com/document/d/1m8b7u97uteHYjXk-4DNglJSpQO8OcZOCzW2tApCNTW4/edit?tab=t.0#heading=h.td5on8bji7s3>

Best,
Andrei

On Fri, Jul 17, 2026 at 8:59 PM Russell Spitzer <[email protected]> wrote:

    Make sure you add it to the dev calendar and send out an
    announcement email :)

    On Fri, Jul 17, 2026 at 12:42 PM Andrei Tserakhau via dev
    <[email protected]> wrote:

        Hi Alexander,

        Yes, I think Aug 5th would be ideal, around 5PM CET  / 8AM PST
        I'll make the calendar slot.

        Best,
        Andrei

        On Fri, Jul 17, 2026 at 3:17 PM Alexander Löser
        <[email protected]> wrote:

            Hi Andrei,

            thanks for bringing up collations in the last community
            sync. If I got it right, the next step would be to gather
            a group of interested folks and set up a dedicated sync -
            preferably with 2+ weeks headsup so that everyone can
            plan accordingly.

            I'm definitely interested in participating in that
            meeting. How about, e.g., Aug 5th or 7th?

            Best,
            Alex

            On 7/15/26 12:06, Alexander Löser wrote:
            Hi Andrei,

            I have some questions/thoughts on the suggestions in
            your latest mail, but I'm happy to defer those for now
            in favor of a more general discussion.

            With regards to your question:
            > how much cross-engine pruning interoperability should
            the format guarantee, versus leave to convention?

            I think this is the right question to ask. In fact, I
            would set the pruning aspect aside for now and only ask:
            How much cross-engine interoperability should the format
            guarantee?

            If I understand your proposal correctly, you were
            suggesting that we should allow each engine to choose
            their ICU version on their own (and skipping pruning if
            need be), rather than pin a specific ICU version at the
            table/schema level.

            I think this suggestion has merit; for example, it will
            make it much easier for engines to upgrade their ICU
            version, independently of Iceberg version changes.

            However, allowing each engine to choose its own ICU
            version has implications beyond pruning - it also
            impacts execution.
            ICU does not guarantee the stability of orderings across
            different versions, i.e., for two strings x and y, x < y
            may hold in version N, but not in version N+1. While
            this usually affects only a small subset of code points,
            ordering changes have occurred with every other ICU
            release for the last couple of years.
            Consequently, the same query may return different
            results on different engines (using different ICU
            versions). This can manifest in various ways - different
            sort orders, more/fewer aggregation results, filtering
            more/less, etc. - which can be very surprising for
            users. A colleague mentioned that at a previous company,
            they had to roll back an ICU library upgrade because
            their users complained about the sort order differences.

            Equality deletes are another complication. Until now,
            whether an equality delete removes a given row is
            unambiguous - every engine agrees. With different ICU
            versions, however, engines may draw different
            conclusions. For example, select * could return
            different rows if two different ICU versions disagree
            whether a string matches one of the deleted values. This
            is, imho, particularly concerning as one DML has the
            potential to cause different results for all subsequent
            queries, even if those queries don't use any collated
            columns. I'm not sure if there would be a way to solve
            this without disallowing equality deletes on collated
            columns.

            Both of the problems described above would disappear if
            we align on one specified ICU version.

            To close the loop, I think the question we need to
            answer is: how much interoperability should the spec
            guarantee?
            I don't have a strong opinion, yet - I mostly want to
            make sure we decide this consciously rather than by
            omission.
            I definitely see value in having consistent results
            across engines.
            At the same time, I'm not sure if we can expect
            consistent results across different engines even today:
            for example, queries involving upper() or lower() may
            produce different results depending on the case mappings
            used by the engine - which depend on the Unicode
            version, too. That said, upper() and lower() are not
            part of the Iceberg spec while collations would be, so
            I'm not sure if this is a good reference point.

            Would be curious to hear what you and the others think.

            Best, Alex


            On 7/3/26 15:56, Andrei Tserakhau via dev wrote:
            Hi Alex,

            Thanks, these are the right questions. Let me answer
            them, but I think all three are really facets of one
            decision worth pulling out, so I'll do that at the end.

            Original values vs sort keys. I don't think the two
            limitations are symmetric. You're right that a bound
            stored under version X may not hold under Y for either
            representation, and a naive reader prunes only on an
            exact version match either way. But they degrade
            differently: a sort key from X is incomparable under Y
            (nothing a Y reader can do with it) while an original
            value is the actual string, so a Y reader can
            re-interpret it. In the case you describe at the end of
            your mail, where only a small code-point range moved
            and a file's values fall outside it, that reader can
            prove the X bound still holds and prune across
            versions. Sort keys foreclose that; original values
            keep it open. So original values are a superset: worst
            case they match sort keys, best case they prune across
            versions.

            Your two sort-key advantages are real, I just don't
            think they belong in the format. Truncation: agreed
            it's hard for collated strings (your abcเก contraction
            case is exactly the trap), so I sidestepped it,
            collation bounds must be tight, a writer that can't
            store the exact min/max omits the bound.
            Collation-aware truncation with
            CollationElementIterator is a possible later
            optimization. Compare cost: pruning is per-file at
            planning time, not per-row, so collator vs byte compare
            is in the noise; and an engine that wants the byte path
            can derive and cache the sort key from the stored
            value. Original values don't block that, they just
            don't bake a version-specific encoding into the format.

            Column vs file-level version. As you say, every engine
            can read regardless, so this is a pruning-performance
            choice, not correctness. In the
            schema-registered-metrics design a file carries bounds
            under a declared (collation, version), and a reader
            prunes any file with a metric for a version it can
            produce, not only files it wrote. So convergence on one
            ICU version gives full cross-engine pruning, same as
            column-level, and a writer or compaction can populate
            several versions at once. It gives the column-level
            benefit by convention without making a version bump a
            format-breaking change.

            [One data point from the engine side, since I'm coming
            at this from the Databricks runtime: our runtime is
            effectively versionless (customers don't pin an ICU
            version, and upgrades happen under them) so "the same
            table read by clients on different ICU versions" isn't
            a corner case for us, it's the default. That's what
            pushes me toward per-file versioning: pinning one
            version per table or column means either forcing the
            whole fleet to upgrade in lockstep or breaking pruning
            on every bump, and neither survives a versionless
            fleet. And in practice most version bumps we've gone
            through don't reorder the data in a given column at
            all, which is exactly why keeping original values, and
            eventually your code-point-range idea, lets a reader
            keep pruning across a bump instead of falling back to a
            scan.]

            Providers. Agreed. I'll tighten the spec to a
            registered set like geo, icu to start, utf8 reserved,
            non-ICU collations added by spec change rather than ad
            hoc. Interop is the whole point and an open namespace
            undercuts it.

            -----------
            Stepping back: I think the three above collapse into
            one question that needs broader alignment than the two
            of us:

            > how much cross-engine pruning interoperability should
            the format guarantee, versus leave to convention?

            Original-vs-sortkey, column-vs-file version, and
            open-vs-restricted providers are all that same tradeoff
            from different angles. It's a values call more than a
            correctness one, and it binds every engine that hasn't
            weighed in yet: Trino, Flink, Spark, PyIceberg, rust.

            From the DBR side I can say the multi-version case is
            real rather than theoretical, but that's one engine's
            vantage point. I'd like to get the interop question in
            front of the other implementers before we fix field
            IDs, the dev list is probably enough for now, and a
            community sync is there if it needs more than async. No
            rush on that; I'd rather let the thread settle the
            mechanics first.

            Best,
            Andrei

            On Fri, Jul 3, 2026 at 12:22 AM Alexander Löser
            <[email protected]> wrote:

                Hi Andrei,

                Thanks for putting together the spec PR and the
                detailed write-up! The approach mostly looks solid
                to me. I have a few questions/initial thoughts
                regarding the changes you proposed (compared to the
                original proposal):

                > 1 - Bounds store original values, not sort keys,
                tagged with a per-file collation version. ICU/CLDR
                sort keys aren't stable across versions, so storing
                keys ties every reader to one exact version;
                original values plus a per-file version (readers
                prune only on an exact match) degrade gracefully
                instead of breaking. The schema keeps the collation
                name unversioned so anyone can read

                If I understand correctly, we’re talking about two
                separate things here:

                 1. Tagging a column vs a single file with a
                    certain ICU version
                 2. Using collation keys vs original strings (the
                    ones that will produce the min/max collation keys)


                For 1, it comes down to a tradeoff:

                  * If we tag the column with the ICU version, we’d
                    force engines to support one agreed-on ICU
                    version if they want to prune files. Engines
                    would be able to prune every file (if they
                    support the specific ICU version), or none at
                    all, so there is more incentive to support a
                    specific version
                  * If we tag individual files with ICU versions,
                    we gain the big advantage that engines do not
                    need to agree on a single ICU version. However,
                    if I understand correctly, this comes at the
                    cost of “fractured” pruning - engines will only
                    be able to prune files that were written by
                    themselves (or rather, with the same ICU
                    version). As a consequence, performance might
                    not really be interoperable between different
                    engines.

                Regardless of the approach we choose, all engines
                should be able to read the data - they might just
                not be able to prune files.

                For 2, I’m not sure if I understand the advantages
                of original strings yet. As you already pointed
                out, the collation keys depend on the ICU version.
                However, if I understand correctly, the same
                limitation would apply to the original strings: the
                sort order may (and does) change between different
                ICU versions, too. As a consequence, we can’t
                assume the original lower/upper bound strings we
                stored for version X will also be lower/upper
                bounds for version Y - at least in the general
                case. So if I understand correctly, we would not
                gain additional pruning opportunities compared to
                using collation keys. Or am I missing something here?

                At the same time, sort keys do have advantages:

                  * Iceberg allows the truncation of upper- and
                    lower bounds. This is trivial for binary
                    collation keys. For original strings, the task
                    becomes significantly harder: truncating at a
                    character boundary, for example, would lead to
                    wrong results, as there are some
                    context-sensitive sequences: e.g., with the
                    CLDR root locale, abcเก < abcเ. I think it
                    might be doable with ICU’s
                    CollationElementIterator, but it will be tricky
                    to get right.
                  * Lower/upper bounds are computed once, but will
                    be compared many times. With original strings,
                    we would need to either convert to the
                    collation key on the fly, or use ICU’s collator
                    for a direct comparison. Both options will be
                    slower than a raw byte-sequence comparison


                There is one scenario where original values would
                shine, though. I analyzed the order-changes between
                various ICU versions: in many cases, only a small
                range of code points changes/is moved. If we had
                additional metadata about which code point ranges a
                file contains (e.g., whether it is ASCII only),
                engines might be able to prove that the original
                string bounds for version X are still valid for
                version Y.
                If I'm not mistaken, this could allow to prune
                across different ICU versions in certain
                situations, which I’d consider a point in favor of
                original values (and file-level ICU versions).



                > 2 - A provider-qualified identifier
                (icu.en_US-ci), leaving room for non-ICU collations
                like Spark's UTF8_LCASE, rather than assuming ICU
                as the sole provider.

                Adding a provider-mechanism sounds like a good
                approach to keep the spec open for future
                collations :slightly_smiling_face: I wonder whether
                we should restrict the set of allowed providers,
                though, similar to how it was done with geo
                
<https://lists.apache.org/thread/r5x0do8f241bpf565rx8s5s3wc9ogp0f>.
                My main motivation for this proposal is
                interoperability. I worry that interoperability
                might suffer or vanish completely if every engine
                can come up with their own definitions.



                Happy to hear your thoughts on this!

                Best, Alex


                On 6/27/26 01:37, Szehon Ho wrote:
                Very nice direction, left some comments on the
                spec proposal.

                Thanks to you folks for working on it !
                Szehon

                On Fri, Jun 26, 2026 at 3:29 AM Andrei Tserakhau
                via dev <[email protected]> wrote:

                    Hi all,

                    I've spend some cycle on the collation
                    discussion and make something more concrete to
                    react to: a spec-change PR plus reference
                    implementations (go and java).

                    - Spec change (apache/iceberg#16972): a
                    "collation" annotation on string fields, and a
                    data_file.collation_bounds field so collated
                    columns stay prunable.
                    - Reference implementation in iceberg-go
                    (apache/iceberg-go#1318): the full path end to
                    end - schema annotation, collation-aware
                    comparison (CLDR/UCA), collation bounds in the
                    manifest, and version-gated data-file pruning,
                    with an Avro round-trip and pruning tests.
                    - A lightweight Java POC (link below): the
                    schema annotation plus a Collator-backed
                    comparator, to match where the discussion is.
                    I deliberately left the manifest/bounds side
                    out of Java for now.

                    The design follows the original proposal but
                    takes a few different turns, mostly to adopt
                    what we learned in Delta. The ones I'd most
                    like input on:

                    1 - Bounds store original values, not sort
                    keys, tagged with a per-file collation
                    version. ICU/CLDR sort keys aren't stable
                    across versions, so storing keys ties every
                    reader to one exact version; original values
                    plus a per-file version (readers prune only on
                    an exact match) degrade gracefully instead of
                    breaking. The schema keeps the collation name
                    unversioned so anyone can read.

                    2 - A provider-qualified identifier
                    (icu.en_US-ci), leaving room for non-ICU
                    collations like Spark's UTF8_LCASE, rather
                    than assuming ICU as the sole provider.

                    3 - One structural question I don't have a
                    strong opinion on yet: I put collation_bounds
                    on data_file as a standalone v3 field, but
                    field id 146 is already the v4 content_stats
                    struct, and collation bounds might belong
                    inside that typed-stats framework instead.
                    Worth settling before we fix field ids.

                    The full set of differences and the
                    reader/writer rules are in the PR description
                    and the write-up. Comments very welcome — both
                    on the calls above and on whether the
                    standalone-field vs content_stats direction is
                    the right one.

                    Best, Andrei

                    - original proposal:
                    
https://docs.google.com/document/d/1m8b7u97uteHYjXk-4DNglJSpQO8OcZOCzW2tApCNTW4/edit?tab=t.0
                    - spec change:
                    https://github.com/apache/iceberg/pull/16972
                    - POC in go:
                    https://github.com/apache/iceberg-go/pull/1318
                    - java POC:
                    
https://github.com/laskoviymishka/iceberg/tree/prototype/collation-support

                    On Mon, Mar 30, 2026 at 10:54 PM Alexander
                    Löser <[email protected]> wrote:

                        Hi Andrei,

                        I'm glad you're interested. Looking
                        forward to collaborate with you!
                        Thanks for all the feedback here and in
                        the doc. I only had a quick glance, but I
                        think you raised some good points. I'll
                        address/respond to your comments as soon
                        as I get the chance, hopefully tomorrow.
                        I think you also left some comments in
                        this mail that are not yet in the doc -
                        I'll move those to a dedicated section at
                        the end of the doc, so we can use the doc
                        as a single source of truth/discussion.

                        > Happy to share our Delta design doc and
                        implementation learnings in more detail.

                        Sure, sounds good :)

                        Best,
                        Alex

                        On 3/29/26 01:25, Andrei Tserakhau via dev
                        wrote:
                        Hi Alexander,

                        This looks really interesting. We've been
                        working on collation support in Delta and
                        have shipped it in production for some
                        time, so this is an area we care about a
                        lot. If this proposal moves forward we'd
                        be happy to collaborate on the design and
                        implementation.

                        The pseudo-field approach for collation
                        metrics is clean and composes well with
                        existing Iceberg infrastructure. The
                        specifier coverage is comprehensive.

                        A few areas worth discussing as this evolves:

                        1 - Sort key stability and versioning

                        ICU sort keys are not stable across
                        versions, so a pinned ICU version bump in
                        a future Iceberg release would invalidate
                        all existing collation metrics. In
                        multi-engine environments, requiring all
                        engines to converge on one ICU version is
                        unrealistic.

                        We store original string values instead
                        of sort keys and allow per-file version
                        annotations -- worth discussing whether
                        something similar could work here.

                        2 - Provider abstraction

                        The proposal assumes ICU as the sole
                        provider, but Spark ships non-ICU
                        collations like UTF8_LCASE that are
                        widely used. A provider or namespace
                        layer would prevent name collisions and
                        support engine-specific collations
                        without future spec changes.

                        3 - Operational surface

                        A few things that turned out
                        correctness-critical in our
                        implementation: partition transforms on
                        collated columns (collation-equal but
                        byte-distinct values in different
                        directories), sort order semantics,
                        equality deletes under collation, and
                        Parquet filter pushdown (must be disabled
                        since Parquet has no collation concept).

                        These don't all need to be solved in v1
                        but would help to scope them.

                        4 - Smaller items (nit's)

                        UTF-8 bounds for the original field id
                        should be "must write" not "should" --
                        otherwise backward compat breaks for
                        non-aware engines. Engine fallback
                        behavior (case-sensitive vs older ICU vs
                        fail) could use a recommended preference
                        order to avoid divergent results across
                        engines. The collation specifier syntax
                        would benefit from a formal grammar.

                        ---

                        Happy to share our Delta design doc and
                        implementation learnings in more detail.
                        Looking forward to the discussion.

                        Best,
                        Andrei

                        On Sat, Mar 28, 2026 at 11:49 PM
                        Alexander Löser <[email protected]>
                        wrote:

                            Hi everyone,

                            this is my first interaction with the
                            Iceberg community, so here a few
                            words about myself:
                            - I'm Alex, a Berlin-based software
                            engineer
                            - I've been working at Snowflake for
                            4 years now
                            - I spend most of my time on data
                            types, particularly binary, strings
                            and collations.

                            I'd like to start a discussion about
                            adding collations to the Iceberg spec.

                            Conceptually, collations are an
                            annotation on the string data type.
                            By default, most engines perform
                            string operations case-sensitively.
                            Collations allow specifying
                            alternative comparison rules. This is
                            useful for achieving, e.g., case- or
                            accent-insensitive string operations,
                            or language-specific string sorting.
                            Collations are supported by many
                            engines: Databricks
                            
<https://docs.databricks.com/aws/en/sql/language-manual/sql-ref-collation>,
                            Spark
                            
<https://spark.apache.org/docs/latest/api/python/reference/pyspark.sql/api/pyspark.sql.functions.collate.html>,
                            Snowflake
                            
<https://docs.snowflake.com/en/sql-reference/collation>,
                            Oracle
                            
<https://docs.oracle.com/en/database/oracle/oracle-database/19/sqlrf/COLLATION.html>
 - to
                            name just a few - this list is not
                            complete.

                            In Snowflake, we see heavy use of the
                            collation feature. Several users have
                            approached us, mentioning they want
                            to migrate to Iceberg tables, but are
                            currently blocked by Iceberg's lack
                            of collation support.

                            Given the widespread support for
                            collations across different engines,
                            I believe introducing collations to
                            Iceberg will increase
                            interoperability and boost its adoption.
                            I'd be curious about your thoughts.

                            *Goal of the proposal*
                            - Support collation specifications
                            for columns
                            - Define how collation bounds should
                            be stored - UTF-8 based bounds are
                            not useful for collated columns

                            *Required Changes*
                            - Extend the schema to let (string)
                            fields be annotated with a collation

                            More details can be found in this doc
                            
<https://docs.google.com/document/d/1m8b7u97uteHYjXk-4DNglJSpQO8OcZOCzW2tApCNTW4/edit?tab=t.0#heading=h.y1ant4w2163k>.

                            I'm also hoping to present the idea
                            in the next community sync.

                            Best, Alex

Reply via email to