+1 (non-binding)
Thanks, that's exactly the distinction I was trying to get at. If the
parametrization is fully captured by the expression in metadata, then we're
aligned: there are no hidden inputs and the expression remains a
well-defined transformation of the row.

- Gianluca

Il giorno gio 3 set 2026 alle ore 21:58 Ryan Blue <[email protected]> ha
scritto:

> Thanks for the background and hilbert details. I understand what you're
> saying, but I don't think it affects this work. However we parameterize
> this function, what we store in metadata needs to completely describe the
> function's operation. We are concerned with functions that are well-defined
> transformations of a single input row. That could be something like
> `hilbert(zvalue(col1, 0, 1), zvalue(col2, 100, 50))` to account for what
> you're talking about.
>
> On Thu, Sep 3, 2026 at 12:17 PM Gianluca Graziadei <
> [email protected]> wrote:
>
>> Hi Ryan,
>>
>> You're right, and apologies: I left the actual case implicit and used
>> hilbert(c1, c2) as shorthand for it. With a fixed per-column bit width the
>> curve index is a pure function of the row, so as written my example doesn't
>> support my point. Let me state the real case explicitly.
>>
>> Hilbert clustering over heterogeneous columns doesn't work on raw values.
>> Each column has a different domain, scale and distribution, and if you feed
>> them to the curve as-is, most of the bits of one dimension end up unused
>> and the clustering degenerates along it.
>> This is not hypothetical: it came out of the discussion with Russell on
>> the Hilbert PR, and it's what the PR 17893 implements on top of the
>> MultiColumn Term refactor (see also this draft
>> https://github.com/GGraziadei/iceberg/pull/1).
>>
>> So in practice each input goes through a per-column transform that maps
>> it onto the available bits, and the expression is: *hilbert(linear(c1),
>> zstdnorm(c2), quantile(c3)) *where linear is a min/max rescaling,
>> zstdnorm a z-score standardization, and quantile a mapping onto quantile
>> boundaries.
>>
>> I don't think this needs a new structure because the proposal already
>> fits it, provided two things are explicit in the spec (RFC 2119 ):
>>
>> - expr MUST be a deterministic function of the referenced field values
>> only; any constant, including fitted parameters, MUST be embedded as a
>> literal in the expression. This is what makes "no hidden inputs" a
>> checkable property rather than an assumption.
>> - Entries in the non-materialized fields table MUST be immutable:
>> changing an expression, including re-fitting its parameters, allocates a
>> new field ID, and IDs MUST NOT be reused (maybe this is implicit).
>>
>> I think this is also where Russell's two points land. A sort order over
>> an expression like the one above is precisely the case where you want to
>> save the computed values, and the lineage question is the transformation
>> re-fit case: under rule 2 the field ID itself carries the lineage, so I
>> don't think a separate ExpressionsID is needed. If instead an entry can be
>> updated in place under the same ID, then we do need versioning. Is
>> immutability the intent?
>>
>> I'll park the CDF suggestion (I will detach in a separate thread that
>> could be useful also for optimizing the Spark physical plan): it's about
>> stats representation, not id allocation however it is important at this
>> level to understand this exigence (=validity scope for expressions) and
>> plan according to it.
>>
>> Cheers,
>> Gianluca
>>
>> Il giorno gio 3 set 2026 alle ore 20:23 Russell Spitzer <
>> [email protected]> ha scritto:
>>
>>> This makes sense for me. I think as we do this we should also define an
>>> output for "Sort Orders" since that's the other main function I think we
>>> have that we would want to save values for and it should look pretty
>>> similar to partition output (just mimicking the input values).)
>>>
>>>
>>> Since these wouldn't live in the "Schema' are we also tracking the
>>> lineage of these expressions over time? Does a table have both a SchemaID
>>> and ExpressionsID?
>>>
>>> On Thu, Sep 3, 2026 at 1:17 PM Ryan Blue <[email protected]> wrote:
>>>
>>>> Thanks for taking a look, Peter and Gianluca.
>>>>
>>>> For the result data type, I was thinking that we would require it if it
>>>> is not a case where we know it can be derived. For instance, we know the
>>>> output type for partition fields so we don't need to specify it and it can
>>>> change. For cases where you may not know the output, like when you have an
>>>> expression, we would require the data type. And I agree that we would
>>>> use bound or id-based references.
>>>>
>>>> I don't think that there is much value in having a special case for
>>>> identity. You already have a field ID for the data, so I don't think there
>>>> is a situation in which you would ever do this.
>>>>
>>>> Also, I see the distinction that Gianluca points out, but this is not
>>>> an issue because there should be no functions in the second category. Maybe
>>>> hilbert was a poor choice for an example. I don't think that we need to
>>>> design for expressions that have hidden inputs.
>>>>
>>>> Ryan
>>>>
>>>> On Tue, Sep 1, 2026 at 4:17 AM Péter Váry <[email protected]>
>>>> wrote:
>>>>
>>>>> Hi Ryan,
>>>>>
>>>>> Thanks for putting this together. This would make index definitions
>>>>> significantly simpler.
>>>>>
>>>>> A couple of thoughts:
>>>>>
>>>>>    - *Schema evolution*: In Iceberg expressions, the result type is
>>>>>    currently defined by the expression itself and may evolve as the table
>>>>>    schema changes (e.g., int -> long). Storing an explicit data-type would
>>>>>    make reading metadata and index data much simpler, but it does remove 
>>>>> some
>>>>>    of that flexibility.
>>>>>    - *FieldId references*: It would probably be a good idea to
>>>>>    require expressions to reference fields by ID (e.g., field(15)) rather 
>>>>> than
>>>>>    by name (col_a) to make them resilient to column renames.
>>>>>    - *Identity expressions*: I expect index definitions to frequently
>>>>>    contain entries like:
>>>>>
>>>>>
>>>>> *              {"field-id": 104, "type": "expr-value", "data-type":
>>>>>    "long", "expr": "<identity(col_a)>"} *
>>>>>    Would it be worth introducing a dedicated type for this case, such
>>>>>    as:
>>>>>
>>>>>
>>>>> *             {"field-id": 104, "type": "identity-value",
>>>>>    "original-field-id": 15} *
>>>>>    It doesn't seem broadly useful outside indexing, but identity
>>>>>    projections in indexes are likely common enough that a more compact
>>>>>    representation may be worth considering.
>>>>>
>>>>> Thanks,
>>>>> Peter
>>>>>
>>>>> Gianluca Graziadei <[email protected]> ezt írta (időpont:
>>>>> 2026. szept. 1., K, 7:02):
>>>>>
>>>>>> Hi Ryan,
>>>>>> I like the proposal; it is concise and clear.
>>>>>>
>>>>>> I would suggest making a clear distinction between two classes of
>>>>>> expressions:
>>>>>>
>>>>>> 1. Expressions that produce deterministic values intrinsically tied
>>>>>> to the input value, e.g. to_lower_case(s).This class is relatively
>>>>>> straightforward to reason about, since the result depends only on the
>>>>>> individual row.
>>>>>> 2. Expressions such as Hilbert that produce values whose meaning
>>>>>> depends on the input value + the whole distribution. Here the main issue 
>>>>>> is
>>>>>> that the underlying distribution matters. If the distribution changes 
>>>>>> over
>>>>>> time, hilbert(c11,c12) computed at time t may no longer be valid at t+1.
>>>>>> This makes the second class harder to handle, because we need to 
>>>>>> establish
>>>>>> for how long a computed value remains valid (=if in the current snapshot
>>>>>> values computed on a previous snapshot are still valid).
>>>>>>
>>>>>> For this second class, I don't think persistence of the computed
>>>>>> value alone is sufficient. We may also need to persist the distribution 
>>>>>> (or
>>>>>> its CDF) against which the value was computed (I am not confident that
>>>>>> min/max can be good enough in this case).
>>>>>>
>>>>>> Before I comment:
>>>>>> 1. Are you targeting to manage these two classes in the same way? It
>>>>>> seems that they have potentially different validity scope.
>>>>>> 2. What about adding a snapshot level CDF struct per column?
>>>>>>
>>>>>> Cheers,
>>>>>> Gianluca
>>>>>>
>>>>>> Il giorno mar 1 set 2026 alle ore 00:25 Ryan Blue <[email protected]>
>>>>>> ha scritto:
>>>>>>
>>>>>>> Hi everyone,
>>>>>>>
>>>>>>> One of the remaining open questions for v4 metadata is how we will
>>>>>>> assign table field IDs for values that are not written into the table. I
>>>>>>> want to propose a solution that I think is going to be flexible, while 
>>>>>>> not
>>>>>>> introducing a lot of churn in the table or REST specs.
>>>>>>>
>>>>>>> Columnar field stats are written into metadata using a simple
>>>>>>> function from table field ID to metadata field ID. We want to reuse 
>>>>>>> what we
>>>>>>> already have working for table fields and keep the spec simple. That 
>>>>>>> means
>>>>>>> we need a way to assign a table field ID to a non-materialized column so
>>>>>>> that we can track its stats. There are a few cases we’ve identified:
>>>>>>>
>>>>>>>    - Partition field output for non-monotonic functions, like 
>>>>>>> bucket(1024,
>>>>>>>    id)
>>>>>>>    - Clustering expressions, like to_lower_case(last_name)
>>>>>>>    - Collation sequence lower and upper bounds
>>>>>>>
>>>>>>> We also discussed a new case this morning in the index sync: we need
>>>>>>> a field ID for a derived value used to organize an index, like 
>>>>>>> hilbert(col_a,
>>>>>>> col_b), because we intend to use table field IDs in index schemas.
>>>>>>>
>>>>>>> Initially, I suggested that we keep a table of expressions and
>>>>>>> assign each one a field ID. But as we started thinking about the use 
>>>>>>> cases
>>>>>>> where we need expressions it became clear that denormalizing *all*
>>>>>>> expressions was adding a lot of complexity for little benefit. For 
>>>>>>> example,
>>>>>>> CHECK constraints won’t have reusable expressions and it makes
>>>>>>> little sense to create them in two parts (expression and constraint).
>>>>>>> Similarly, it is awkward to model a collation sequence as an expression,
>>>>>>> and we don’t need to rebuild partition specs just to assign field IDs.
>>>>>>> However, we also don’t want to just embed table field IDs in every one 
>>>>>>> of
>>>>>>> these structures.
>>>>>>>
>>>>>>> My proposal is to directly model what we want: one table of fields
>>>>>>> that are not materialized in the table, but are assigned IDs for stats 
>>>>>>> or
>>>>>>> other purposes. This would take a few forms:
>>>>>>>
>>>>>>>    - Partition output value: {"field-id": 102, "type":
>>>>>>>    "partition-value", "partition-field-id": 1000}
>>>>>>>    - Collation sequence: {"field-id": 103, "type":
>>>>>>>    "collation-bounds", "collation-seq-id": 1}
>>>>>>>    - Value expression results: {"field-id": 104, "type":
>>>>>>>    "expr-value", "data-type": "long", "expr": <hilbert(col_a, col_b) 
>>>>>>> expr>}
>>>>>>>
>>>>>>> This representation leaves existing structures alone and is a single
>>>>>>> place outside of schema to allocate table field IDs. This can be 
>>>>>>> expanded
>>>>>>> with new types later when we want to add new structures, like a 
>>>>>>> cluster-by
>>>>>>> spec.
>>>>>>>
>>>>>>> I think this is a fairly clean way to move forward and solve two
>>>>>>> challenges that we’re currently hitting. We'll discuss this in the next 
>>>>>>> v4
>>>>>>> metadata sync, but in the meantime please reply with feedback if you 
>>>>>>> have
>>>>>>> an opinion.
>>>>>>>
>>>>>>> Thanks,
>>>>>>>
>>>>>>> Ryan
>>>>>>>
>>>>>>

Reply via email to