Hi Manu,

Good question.

This is not for single-value lookups. `Table.Labels()` already covers that.
The metadata table is for consumers that work with metadata as Arrow.

The design is the same across clients, but PyIceberg is probably the
clearest example. In PyIceberg, inspect already returns PyArrow, so labels
fit naturally there. You can do something like
`pa.concat_tables([t.inspect.labels() for t in tables])` and then
filter/group across many tables for things like PII, ownership, or cost
attribution. There is no DESCRIBE-style surface in the client.

For Go/Rust/C++ the main benefit is consistency. iceberg-go already exposes
snapshots, history, manifests, etc. through InspectTable; labels can use
the same path instead of special-casing the nested map. I also plan to add
`iceberg inspect <table> labels` to the CLI.

One caveat: only REST catalogs return labels today, so this is empty
elsewhere, and for simple per-table access the typed accessor is enough.

That’s why I think this fits the client inspect APIs better than
Java/Spark, where the SQL surface already exists.

Thanks,
Andrei

On Sat, Oct 3, 2026 at 6:22 PM Manu Zhang <[email protected]> wrote:

> Hi Andrei,
>
> Could you please elaborate on what issues you are trying to solve in
> iceberg-go with a labels metadata table?
>
> Andrei Tserakhau via dev <[email protected]>于2026年10月3日 周六07:17写道:
>
>> Hi all,
>>
>> Reviving this with a narrower question after the same issue came up again
>> in iceberg-go [1].
>>
>> The java labels metadata-table PR [0] was closed because metadata tables
>> in Java/Spark are expected to be deterministic projections of table
>> metadata. I think that still makes sense for the engine SQL / Spark surface.
>>
>> The non-Java clients are a bit different, though. In PyIceberg,
>> iceberg-rust, and iceberg-go, inspect is a library read API, not a SQL
>> metadata table, and there is no convinient alternative.
>>
>> My suggestion is to keep the Java/Spark decision as-is, but let each
>> client expose catalog-provided data through its inspect API if useful. For
>> labels, that means documenting it as catalog-provided, captured at load
>> time, and empty when the catalog returns none.
>>
>> If that split sounds reasonable, I'll proceed with the iceberg-go PR [1]
>> and mirror it in rust and python.
>>
>> Any objections to treating the client inspect API separately from the
>> engine metadata-table contract?
>>
>> Thanks,
>> Andrei
>>
>> [0] https://github.com/apache/iceberg/pull/18048 (Java core labels
>> metadata table, closed)
>> [1] https://github.com/apache/iceberg-go/pull/2101 (iceberg-go labels
>> inspect table)
>>
>> On Tue, Sep 29, 2026 at 12:19 AM Andrei Tserakhau <
>> [email protected]> wrote:
>>
>>> Fair - I think this brings the focus back to the main question: why do
>>> we need
>>> SQL at all.
>>>
>>> My answer is that there's a class of consumers whose only interface is
>>> SQL -
>>> SQL-native discovery / governance tooling, and most notably LLM agents
>>> that
>>> explore a warehouse through SQL. For them a Java capability
>>> (SupportsLabels) is
>>> not a thing; they can only reason over what they can query.
>>>
>>> What can cover that is `DESCRIBE` - #18049 surfaces both object and
>>> field labels
>>> there (labels.object.*, labels.field.<id>.*). So if there's no urge to
>>> join
>>> stuff, I see the point - we can keep it simple and not do the metadata
>>> table.
>>>
>>> We can revisit this topic later, once we have more data points and use
>>> cases,
>>> but for now I agree - it's not needed.
>>>
>>> Best,
>>> Andrei
>>>
>>> On Mon, Sep 28, 2026 at 11:30 PM Ryan Blue <[email protected]> wrote:
>>>
>>>> I don't agree with the composition argument. The example query you
>>>> provided doesn't make sense because there is no ON clause so you end up
>>>> with a cartesian join. Luckily, since you're looking for a specific label,
>>>> "owner", you end up with just one label and will aggregate the entire set
>>>> of files. So you end up with a result that looks reasonable, but you're
>>>> really just running two unrelated queries here: one to aggregate the total
>>>> size of live files in the table, and one to select the owner.
>>>>
>>>> I think that means that the only use case here is to expose this data
>>>> to users. But I think that this reasoning is that we need a metadata table
>>>> because we need SQL interaction and we need SQL interaction because . . . ?
>>>> It's a nice-to-have, sure, but I'm not convinced that anyone would miss it
>>>> if we didn't expose this directly to users.
>>>>
>>>> On Fri, Sep 25, 2026 at 2:40 PM Andrei Tserakhau via dev <
>>>> [email protected]> wrote:
>>>>
>>>>> Hi Ryan,
>>>>>
>>>>> Agree here: the engine-facing consumption (cost attribution, policy
>>>>> attachment) goes through SupportsLabels, no table needed. The table is for
>>>>> the other consumer - SQL/people or AI Agent :). Some cases that i see 
>>>>> here:
>>>>>
>>>>> 1) Exploration. "which columns are classified as X here" is a query a
>>>>> person runs, not something an engine surfaces:
>>>>>
>>>>>     SELECT field_name, key, value
>>>>>     FROM prod.db.orders.labels
>>>>>     WHERE scope = 'field' AND key = 'classification';
>>>>>
>>>>> 2) Composition. .labels joins with .files / .partitions / .snapshots
>>>>> in one query - e.g. attribute bytes to an owner label, i.e. cost
>>>>> attribution as a report someone runs, not a log an engine emits:
>>>>>
>>>>>     SELECT l.value AS owner, SUM(f.file_size_in_bytes) AS bytes
>>>>>     FROM prod.db.orders.files f, prod.db.orders.labels l
>>>>>     WHERE l.scope = 'object' AND l.key = 'owner'
>>>>>     GROUP BY l.value;
>>>>>
>>>>> I think the key value to have a metadata table is joinability, you
>>>>> can’t have it with a programmatic label API.
>>>>>
>>>>> So there is a place for the SQL consumer, the engine path is a
>>>>> different usecase and the co-live together.
>>>>>
>>>>> Thanks,
>>>>> Andrei
>>>>>
>>>>>
>>>>> On Fri, Sep 25, 2026 at 10:51 PM Ryan Blue <[email protected]> wrote:
>>>>>
>>>>>> > Are we OK with catalog-provided metadata tables as a separate
>>>>>> category?
>>>>>>
>>>>>> I'm okay with providing metadata through a system table like this, as
>>>>>> long as we think that people will want to access this data that way.
>>>>>> Question 3, "If no, what should the SQL surface for labels be instead?"
>>>>>> makes me think that a SQL surface is _assumed_ to be needed.
>>>>>>
>>>>>> I don't think it is necessarily the case that we need to expose these
>>>>>> for SQL users. I thought that we wanted labels to expose additional 
>>>>>> context
>>>>>> to engines for things like cost attribution logs or attaching an engine's
>>>>>> policy to a table. That doesn't require a table-like user surface.
>>>>>>
>>>>>> I'm fine adding a metadata table if there's a use for it, but if we
>>>>>> don't need one then it's simpler not to add and maintain it. And that
>>>>>> avoids needing to answer questions like this as well.
>>>>>>
>>>>>>
>>>>>>
>>>>>> On Fri, Sep 25, 2026 at 1:13 PM Andrei Tserakhau via dev <
>>>>>> [email protected]> wrote:
>>>>>>
>>>>>>> Hi all,
>>>>>>>
>>>>>>> Labels in the REST spec recently landed [1]. A catalog can now
>>>>>>> expose object-level and per-field labels on load-table responses. As 
>>>>>>> one of
>>>>>>> the follow-ups, there is a proposal to add a .labels metadata table
>>>>>>> backed by catalog data.
>>>>>>>
>>>>>>> During discussion of the follow-ups ([3], [4]), Peter raised a good
>>>>>>> question on [3]: every metadata table today is derived from table 
>>>>>>> metadata.
>>>>>>> .labels would be different because the data comes from the catalog,
>>>>>>> may vary by catalog, and may be absent if the catalog has nothing to 
>>>>>>> return.
>>>>>>>
>>>>>>> I think this is less about Labels itself and more about a new *kind*
>>>>>>> of data we would expose: metadata owned by the catalog rather than by
>>>>>>> storage. Labels would be the first example, but later the same pattern
>>>>>>> could be used to expose other catalog information as metadata tables.
>>>>>>>
>>>>>>> So the broader question is: do we want metadata tables to also
>>>>>>> expose catalog-provided information?
>>>>>>>
>>>>>>> I think labels are a reasonable first case. They are structured,
>>>>>>> useful to query via SQL, and fit the same access pattern as
>>>>>>> .snapshots or .partitions. The table can stay read-only, limited to
>>>>>>> the spec-defined shape, and empty when the catalog returns nothing.
>>>>>>>
>>>>>>> The tradeoff is that this breaks the current assumption that
>>>>>>> metadata tables are deterministic projections of table metadata. It’s 
>>>>>>> not
>>>>>>> explicitly written anywhere, but that assumption exists today.
>>>>>>>
>>>>>>> To summarize the questions:
>>>>>>>
>>>>>>>    1.
>>>>>>>
>>>>>>>    Are we OK with catalog-provided metadata tables as a separate
>>>>>>>    category?
>>>>>>>    2.
>>>>>>>
>>>>>>>    If yes, should we mark them somehow so they are clearly
>>>>>>>    different from spec-backed metadata tables?
>>>>>>>    3.
>>>>>>>
>>>>>>>    If no, what should the SQL surface for labels be instead?
>>>>>>>
>>>>>>> Thanks,
>>>>>>>
>>>>>>> Andrei
>>>>>>>
>>>>>>> [1] REST spec labels: https://github.com/apache/iceberg/pull/15750
>>>>>>> [2] SupportsLabels: https://github.com/apache/iceberg/pull/18046
>>>>>>> [3] Labels metadata table:
>>>>>>> https://github.com/apache/iceberg/pull/18048
>>>>>>> [4] Spark DESCRIBE: https://github.com/apache/iceberg/pull/18049
>>>>>>> [5] Design:
>>>>>>> https://docs.google.com/document/d/1aj-6JlfBiMYEEVtNuh5WLMOrRQiMCcyYUGbouPM4hXI/edit?tab=t.0#heading=h.2w0kmp1v1gwv
>>>>>>>
>>>>>>

Reply via email to