Hi all,

Reviving this with a narrower question after the same issue came up again
in iceberg-go [1].

The java labels metadata-table PR [0] was closed because metadata tables in
Java/Spark are expected to be deterministic projections of table metadata.
I think that still makes sense for the engine SQL / Spark surface.

The non-Java clients are a bit different, though. In PyIceberg,
iceberg-rust, and iceberg-go, inspect is a library read API, not a SQL
metadata table, and there is no convinient alternative.

My suggestion is to keep the Java/Spark decision as-is, but let each client
expose catalog-provided data through its inspect API if useful. For labels,
that means documenting it as catalog-provided, captured at load time, and
empty when the catalog returns none.

If that split sounds reasonable, I'll proceed with the iceberg-go PR [1]
and mirror it in rust and python.

Any objections to treating the client inspect API separately from the
engine metadata-table contract?

Thanks,
Andrei

[0] https://github.com/apache/iceberg/pull/18048 (Java core labels metadata
table, closed)
[1] https://github.com/apache/iceberg-go/pull/2101 (iceberg-go labels
inspect table)

On Tue, Sep 29, 2026 at 12:19 AM Andrei Tserakhau <
[email protected]> wrote:

> Fair - I think this brings the focus back to the main question: why do we
> need
> SQL at all.
>
> My answer is that there's a class of consumers whose only interface is SQL
> -
> SQL-native discovery / governance tooling, and most notably LLM agents that
> explore a warehouse through SQL. For them a Java capability
> (SupportsLabels) is
> not a thing; they can only reason over what they can query.
>
> What can cover that is `DESCRIBE` - #18049 surfaces both object and field
> labels
> there (labels.object.*, labels.field.<id>.*). So if there's no urge to join
> stuff, I see the point - we can keep it simple and not do the metadata
> table.
>
> We can revisit this topic later, once we have more data points and use
> cases,
> but for now I agree - it's not needed.
>
> Best,
> Andrei
>
> On Mon, Sep 28, 2026 at 11:30 PM Ryan Blue <[email protected]> wrote:
>
>> I don't agree with the composition argument. The example query you
>> provided doesn't make sense because there is no ON clause so you end up
>> with a cartesian join. Luckily, since you're looking for a specific label,
>> "owner", you end up with just one label and will aggregate the entire set
>> of files. So you end up with a result that looks reasonable, but you're
>> really just running two unrelated queries here: one to aggregate the total
>> size of live files in the table, and one to select the owner.
>>
>> I think that means that the only use case here is to expose this data to
>> users. But I think that this reasoning is that we need a metadata table
>> because we need SQL interaction and we need SQL interaction because . . . ?
>> It's a nice-to-have, sure, but I'm not convinced that anyone would miss it
>> if we didn't expose this directly to users.
>>
>> On Fri, Sep 25, 2026 at 2:40 PM Andrei Tserakhau via dev <
>> [email protected]> wrote:
>>
>>> Hi Ryan,
>>>
>>> Agree here: the engine-facing consumption (cost attribution, policy
>>> attachment) goes through SupportsLabels, no table needed. The table is for
>>> the other consumer - SQL/people or AI Agent :). Some cases that i see here:
>>>
>>> 1) Exploration. "which columns are classified as X here" is a query a
>>> person runs, not something an engine surfaces:
>>>
>>>     SELECT field_name, key, value
>>>     FROM prod.db.orders.labels
>>>     WHERE scope = 'field' AND key = 'classification';
>>>
>>> 2) Composition. .labels joins with .files / .partitions / .snapshots in
>>> one query - e.g. attribute bytes to an owner label, i.e. cost attribution
>>> as a report someone runs, not a log an engine emits:
>>>
>>>     SELECT l.value AS owner, SUM(f.file_size_in_bytes) AS bytes
>>>     FROM prod.db.orders.files f, prod.db.orders.labels l
>>>     WHERE l.scope = 'object' AND l.key = 'owner'
>>>     GROUP BY l.value;
>>>
>>> I think the key value to have a metadata table is joinability, you can’t
>>> have it with a programmatic label API.
>>>
>>> So there is a place for the SQL consumer, the engine path is a different
>>> usecase and the co-live together.
>>>
>>> Thanks,
>>> Andrei
>>>
>>>
>>> On Fri, Sep 25, 2026 at 10:51 PM Ryan Blue <[email protected]> wrote:
>>>
>>>> > Are we OK with catalog-provided metadata tables as a separate
>>>> category?
>>>>
>>>> I'm okay with providing metadata through a system table like this, as
>>>> long as we think that people will want to access this data that way.
>>>> Question 3, "If no, what should the SQL surface for labels be instead?"
>>>> makes me think that a SQL surface is _assumed_ to be needed.
>>>>
>>>> I don't think it is necessarily the case that we need to expose these
>>>> for SQL users. I thought that we wanted labels to expose additional context
>>>> to engines for things like cost attribution logs or attaching an engine's
>>>> policy to a table. That doesn't require a table-like user surface.
>>>>
>>>> I'm fine adding a metadata table if there's a use for it, but if we
>>>> don't need one then it's simpler not to add and maintain it. And that
>>>> avoids needing to answer questions like this as well.
>>>>
>>>>
>>>>
>>>> On Fri, Sep 25, 2026 at 1:13 PM Andrei Tserakhau via dev <
>>>> [email protected]> wrote:
>>>>
>>>>> Hi all,
>>>>>
>>>>> Labels in the REST spec recently landed [1]. A catalog can now expose
>>>>> object-level and per-field labels on load-table responses. As one of the
>>>>> follow-ups, there is a proposal to add a .labels metadata table
>>>>> backed by catalog data.
>>>>>
>>>>> During discussion of the follow-ups ([3], [4]), Peter raised a good
>>>>> question on [3]: every metadata table today is derived from table 
>>>>> metadata.
>>>>> .labels would be different because the data comes from the catalog,
>>>>> may vary by catalog, and may be absent if the catalog has nothing to 
>>>>> return.
>>>>>
>>>>> I think this is less about Labels itself and more about a new *kind*
>>>>> of data we would expose: metadata owned by the catalog rather than by
>>>>> storage. Labels would be the first example, but later the same pattern
>>>>> could be used to expose other catalog information as metadata tables.
>>>>>
>>>>> So the broader question is: do we want metadata tables to also expose
>>>>> catalog-provided information?
>>>>>
>>>>> I think labels are a reasonable first case. They are structured,
>>>>> useful to query via SQL, and fit the same access pattern as .snapshots
>>>>> or .partitions. The table can stay read-only, limited to the
>>>>> spec-defined shape, and empty when the catalog returns nothing.
>>>>>
>>>>> The tradeoff is that this breaks the current assumption that metadata
>>>>> tables are deterministic projections of table metadata. It’s not 
>>>>> explicitly
>>>>> written anywhere, but that assumption exists today.
>>>>>
>>>>> To summarize the questions:
>>>>>
>>>>>    1.
>>>>>
>>>>>    Are we OK with catalog-provided metadata tables as a separate
>>>>>    category?
>>>>>    2.
>>>>>
>>>>>    If yes, should we mark them somehow so they are clearly different
>>>>>    from spec-backed metadata tables?
>>>>>    3.
>>>>>
>>>>>    If no, what should the SQL surface for labels be instead?
>>>>>
>>>>> Thanks,
>>>>>
>>>>> Andrei
>>>>>
>>>>> [1] REST spec labels: https://github.com/apache/iceberg/pull/15750
>>>>> [2] SupportsLabels: https://github.com/apache/iceberg/pull/18046
>>>>> [3] Labels metadata table:
>>>>> https://github.com/apache/iceberg/pull/18048
>>>>> [4] Spark DESCRIBE: https://github.com/apache/iceberg/pull/18049
>>>>> [5] Design:
>>>>> https://docs.google.com/document/d/1aj-6JlfBiMYEEVtNuh5WLMOrRQiMCcyYUGbouPM4hXI/edit?tab=t.0#heading=h.2w0kmp1v1gwv
>>>>>
>>>>

Reply via email to