Hi Andrei,

Thanks for the clarification. However, the API doesn't look right to me
because we can't distinguish whether a table lacks a label or if
the catalog doesn't support it.

Thanks,
Manu

On Sun, Oct 4, 2026 at 2:35 AM Andrei Tserakhau <
[email protected]> wrote:

> Hi Manu,
>
> Good question.
>
> This is not for single-value lookups. `Table.Labels()` already covers
> that. The metadata table is for consumers that work with metadata as Arrow.
>
> The design is the same across clients, but PyIceberg is probably the
> clearest example. In PyIceberg, inspect already returns PyArrow, so labels
> fit naturally there. You can do something like
> `pa.concat_tables([t.inspect.labels() for t in tables])` and then
> filter/group across many tables for things like PII, ownership, or cost
> attribution. There is no DESCRIBE-style surface in the client.
>
> For Go/Rust/C++ the main benefit is consistency. iceberg-go already
> exposes snapshots, history, manifests, etc. through InspectTable; labels
> can use the same path instead of special-casing the nested map. I also plan
> to add `iceberg inspect <table> labels` to the CLI.
>
> One caveat: only REST catalogs return labels today, so this is empty
> elsewhere, and for simple per-table access the typed accessor is enough.
>
> That’s why I think this fits the client inspect APIs better than
> Java/Spark, where the SQL surface already exists.
>
> Thanks,
> Andrei
>
> On Sat, Oct 3, 2026 at 6:22 PM Manu Zhang <[email protected]> wrote:
>
>> Hi Andrei,
>>
>> Could you please elaborate on what issues you are trying to solve in
>> iceberg-go with a labels metadata table?
>>
>> Andrei Tserakhau via dev <[email protected]>于2026年10月3日 周六07:17写道:
>>
>>> Hi all,
>>>
>>> Reviving this with a narrower question after the same issue came up
>>> again in iceberg-go [1].
>>>
>>> The java labels metadata-table PR [0] was closed because metadata tables
>>> in Java/Spark are expected to be deterministic projections of table
>>> metadata. I think that still makes sense for the engine SQL / Spark surface.
>>>
>>> The non-Java clients are a bit different, though. In PyIceberg,
>>> iceberg-rust, and iceberg-go, inspect is a library read API, not a SQL
>>> metadata table, and there is no convinient alternative.
>>>
>>> My suggestion is to keep the Java/Spark decision as-is, but let each
>>> client expose catalog-provided data through its inspect API if useful. For
>>> labels, that means documenting it as catalog-provided, captured at load
>>> time, and empty when the catalog returns none.
>>>
>>> If that split sounds reasonable, I'll proceed with the iceberg-go PR [1]
>>> and mirror it in rust and python.
>>>
>>> Any objections to treating the client inspect API separately from the
>>> engine metadata-table contract?
>>>
>>> Thanks,
>>> Andrei
>>>
>>> [0] https://github.com/apache/iceberg/pull/18048 (Java core labels
>>> metadata table, closed)
>>> [1] https://github.com/apache/iceberg-go/pull/2101 (iceberg-go labels
>>> inspect table)
>>>
>>> On Tue, Sep 29, 2026 at 12:19 AM Andrei Tserakhau <
>>> [email protected]> wrote:
>>>
>>>> Fair - I think this brings the focus back to the main question: why do
>>>> we need
>>>> SQL at all.
>>>>
>>>> My answer is that there's a class of consumers whose only interface is
>>>> SQL -
>>>> SQL-native discovery / governance tooling, and most notably LLM agents
>>>> that
>>>> explore a warehouse through SQL. For them a Java capability
>>>> (SupportsLabels) is
>>>> not a thing; they can only reason over what they can query.
>>>>
>>>> What can cover that is `DESCRIBE` - #18049 surfaces both object and
>>>> field labels
>>>> there (labels.object.*, labels.field.<id>.*). So if there's no urge to
>>>> join
>>>> stuff, I see the point - we can keep it simple and not do the metadata
>>>> table.
>>>>
>>>> We can revisit this topic later, once we have more data points and use
>>>> cases,
>>>> but for now I agree - it's not needed.
>>>>
>>>> Best,
>>>> Andrei
>>>>
>>>> On Mon, Sep 28, 2026 at 11:30 PM Ryan Blue <[email protected]> wrote:
>>>>
>>>>> I don't agree with the composition argument. The example query you
>>>>> provided doesn't make sense because there is no ON clause so you end up
>>>>> with a cartesian join. Luckily, since you're looking for a specific label,
>>>>> "owner", you end up with just one label and will aggregate the entire set
>>>>> of files. So you end up with a result that looks reasonable, but you're
>>>>> really just running two unrelated queries here: one to aggregate the total
>>>>> size of live files in the table, and one to select the owner.
>>>>>
>>>>> I think that means that the only use case here is to expose this data
>>>>> to users. But I think that this reasoning is that we need a metadata table
>>>>> because we need SQL interaction and we need SQL interaction because . . . 
>>>>> ?
>>>>> It's a nice-to-have, sure, but I'm not convinced that anyone would miss it
>>>>> if we didn't expose this directly to users.
>>>>>
>>>>> On Fri, Sep 25, 2026 at 2:40 PM Andrei Tserakhau via dev <
>>>>> [email protected]> wrote:
>>>>>
>>>>>> Hi Ryan,
>>>>>>
>>>>>> Agree here: the engine-facing consumption (cost attribution, policy
>>>>>> attachment) goes through SupportsLabels, no table needed. The table is 
>>>>>> for
>>>>>> the other consumer - SQL/people or AI Agent :). Some cases that i see 
>>>>>> here:
>>>>>>
>>>>>> 1) Exploration. "which columns are classified as X here" is a query a
>>>>>> person runs, not something an engine surfaces:
>>>>>>
>>>>>>     SELECT field_name, key, value
>>>>>>     FROM prod.db.orders.labels
>>>>>>     WHERE scope = 'field' AND key = 'classification';
>>>>>>
>>>>>> 2) Composition. .labels joins with .files / .partitions / .snapshots
>>>>>> in one query - e.g. attribute bytes to an owner label, i.e. cost
>>>>>> attribution as a report someone runs, not a log an engine emits:
>>>>>>
>>>>>>     SELECT l.value AS owner, SUM(f.file_size_in_bytes) AS bytes
>>>>>>     FROM prod.db.orders.files f, prod.db.orders.labels l
>>>>>>     WHERE l.scope = 'object' AND l.key = 'owner'
>>>>>>     GROUP BY l.value;
>>>>>>
>>>>>> I think the key value to have a metadata table is joinability, you
>>>>>> can’t have it with a programmatic label API.
>>>>>>
>>>>>> So there is a place for the SQL consumer, the engine path is a
>>>>>> different usecase and the co-live together.
>>>>>>
>>>>>> Thanks,
>>>>>> Andrei
>>>>>>
>>>>>>
>>>>>> On Fri, Sep 25, 2026 at 10:51 PM Ryan Blue <[email protected]> wrote:
>>>>>>
>>>>>>> > Are we OK with catalog-provided metadata tables as a separate
>>>>>>> category?
>>>>>>>
>>>>>>> I'm okay with providing metadata through a system table like this,
>>>>>>> as long as we think that people will want to access this data that way.
>>>>>>> Question 3, "If no, what should the SQL surface for labels be instead?"
>>>>>>> makes me think that a SQL surface is _assumed_ to be needed.
>>>>>>>
>>>>>>> I don't think it is necessarily the case that we need to expose
>>>>>>> these for SQL users. I thought that we wanted labels to expose
>>>>>>> additional context to engines for things like cost attribution logs or
>>>>>>> attaching an engine's policy to a table. That doesn't require a 
>>>>>>> table-like
>>>>>>> user surface.
>>>>>>>
>>>>>>> I'm fine adding a metadata table if there's a use for it, but if we
>>>>>>> don't need one then it's simpler not to add and maintain it. And that
>>>>>>> avoids needing to answer questions like this as well.
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>> On Fri, Sep 25, 2026 at 1:13 PM Andrei Tserakhau via dev <
>>>>>>> [email protected]> wrote:
>>>>>>>
>>>>>>>> Hi all,
>>>>>>>>
>>>>>>>> Labels in the REST spec recently landed [1]. A catalog can now
>>>>>>>> expose object-level and per-field labels on load-table responses. As 
>>>>>>>> one of
>>>>>>>> the follow-ups, there is a proposal to add a .labels metadata
>>>>>>>> table backed by catalog data.
>>>>>>>>
>>>>>>>> During discussion of the follow-ups ([3], [4]), Peter raised a good
>>>>>>>> question on [3]: every metadata table today is derived from table 
>>>>>>>> metadata.
>>>>>>>> .labels would be different because the data comes from the
>>>>>>>> catalog, may vary by catalog, and may be absent if the catalog has 
>>>>>>>> nothing
>>>>>>>> to return.
>>>>>>>>
>>>>>>>> I think this is less about Labels itself and more about a new
>>>>>>>> *kind* of data we would expose: metadata owned by the catalog
>>>>>>>> rather than by storage. Labels would be the first example, but later 
>>>>>>>> the
>>>>>>>> same pattern could be used to expose other catalog information as 
>>>>>>>> metadata
>>>>>>>> tables.
>>>>>>>>
>>>>>>>> So the broader question is: do we want metadata tables to also
>>>>>>>> expose catalog-provided information?
>>>>>>>>
>>>>>>>> I think labels are a reasonable first case. They are structured,
>>>>>>>> useful to query via SQL, and fit the same access pattern as
>>>>>>>> .snapshots or .partitions. The table can stay read-only, limited
>>>>>>>> to the spec-defined shape, and empty when the catalog returns nothing.
>>>>>>>>
>>>>>>>> The tradeoff is that this breaks the current assumption that
>>>>>>>> metadata tables are deterministic projections of table metadata. It’s 
>>>>>>>> not
>>>>>>>> explicitly written anywhere, but that assumption exists today.
>>>>>>>>
>>>>>>>> To summarize the questions:
>>>>>>>>
>>>>>>>>    1.
>>>>>>>>
>>>>>>>>    Are we OK with catalog-provided metadata tables as a separate
>>>>>>>>    category?
>>>>>>>>    2.
>>>>>>>>
>>>>>>>>    If yes, should we mark them somehow so they are clearly
>>>>>>>>    different from spec-backed metadata tables?
>>>>>>>>    3.
>>>>>>>>
>>>>>>>>    If no, what should the SQL surface for labels be instead?
>>>>>>>>
>>>>>>>> Thanks,
>>>>>>>>
>>>>>>>> Andrei
>>>>>>>>
>>>>>>>> [1] REST spec labels: https://github.com/apache/iceberg/pull/15750
>>>>>>>> [2] SupportsLabels: https://github.com/apache/iceberg/pull/18046
>>>>>>>> [3] Labels metadata table:
>>>>>>>> https://github.com/apache/iceberg/pull/18048
>>>>>>>> [4] Spark DESCRIBE: https://github.com/apache/iceberg/pull/18049
>>>>>>>> [5] Design:
>>>>>>>> https://docs.google.com/document/d/1aj-6JlfBiMYEEVtNuh5WLMOrRQiMCcyYUGbouPM4hXI/edit?tab=t.0#heading=h.2w0kmp1v1gwv
>>>>>>>>
>>>>>>>

Reply via email to