Hi Andrei, Thanks for the clarification. However, the API doesn't look right to me because we can't distinguish whether a table lacks a label or if the catalog doesn't support it.
Thanks, Manu On Sun, Oct 4, 2026 at 2:35 AM Andrei Tserakhau < [email protected]> wrote: > Hi Manu, > > Good question. > > This is not for single-value lookups. `Table.Labels()` already covers > that. The metadata table is for consumers that work with metadata as Arrow. > > The design is the same across clients, but PyIceberg is probably the > clearest example. In PyIceberg, inspect already returns PyArrow, so labels > fit naturally there. You can do something like > `pa.concat_tables([t.inspect.labels() for t in tables])` and then > filter/group across many tables for things like PII, ownership, or cost > attribution. There is no DESCRIBE-style surface in the client. > > For Go/Rust/C++ the main benefit is consistency. iceberg-go already > exposes snapshots, history, manifests, etc. through InspectTable; labels > can use the same path instead of special-casing the nested map. I also plan > to add `iceberg inspect <table> labels` to the CLI. > > One caveat: only REST catalogs return labels today, so this is empty > elsewhere, and for simple per-table access the typed accessor is enough. > > That’s why I think this fits the client inspect APIs better than > Java/Spark, where the SQL surface already exists. > > Thanks, > Andrei > > On Sat, Oct 3, 2026 at 6:22 PM Manu Zhang <[email protected]> wrote: > >> Hi Andrei, >> >> Could you please elaborate on what issues you are trying to solve in >> iceberg-go with a labels metadata table? >> >> Andrei Tserakhau via dev <[email protected]>于2026年10月3日 周六07:17写道: >> >>> Hi all, >>> >>> Reviving this with a narrower question after the same issue came up >>> again in iceberg-go [1]. >>> >>> The java labels metadata-table PR [0] was closed because metadata tables >>> in Java/Spark are expected to be deterministic projections of table >>> metadata. I think that still makes sense for the engine SQL / Spark surface. >>> >>> The non-Java clients are a bit different, though. In PyIceberg, >>> iceberg-rust, and iceberg-go, inspect is a library read API, not a SQL >>> metadata table, and there is no convinient alternative. >>> >>> My suggestion is to keep the Java/Spark decision as-is, but let each >>> client expose catalog-provided data through its inspect API if useful. For >>> labels, that means documenting it as catalog-provided, captured at load >>> time, and empty when the catalog returns none. >>> >>> If that split sounds reasonable, I'll proceed with the iceberg-go PR [1] >>> and mirror it in rust and python. >>> >>> Any objections to treating the client inspect API separately from the >>> engine metadata-table contract? >>> >>> Thanks, >>> Andrei >>> >>> [0] https://github.com/apache/iceberg/pull/18048 (Java core labels >>> metadata table, closed) >>> [1] https://github.com/apache/iceberg-go/pull/2101 (iceberg-go labels >>> inspect table) >>> >>> On Tue, Sep 29, 2026 at 12:19 AM Andrei Tserakhau < >>> [email protected]> wrote: >>> >>>> Fair - I think this brings the focus back to the main question: why do >>>> we need >>>> SQL at all. >>>> >>>> My answer is that there's a class of consumers whose only interface is >>>> SQL - >>>> SQL-native discovery / governance tooling, and most notably LLM agents >>>> that >>>> explore a warehouse through SQL. For them a Java capability >>>> (SupportsLabels) is >>>> not a thing; they can only reason over what they can query. >>>> >>>> What can cover that is `DESCRIBE` - #18049 surfaces both object and >>>> field labels >>>> there (labels.object.*, labels.field.<id>.*). So if there's no urge to >>>> join >>>> stuff, I see the point - we can keep it simple and not do the metadata >>>> table. >>>> >>>> We can revisit this topic later, once we have more data points and use >>>> cases, >>>> but for now I agree - it's not needed. >>>> >>>> Best, >>>> Andrei >>>> >>>> On Mon, Sep 28, 2026 at 11:30 PM Ryan Blue <[email protected]> wrote: >>>> >>>>> I don't agree with the composition argument. The example query you >>>>> provided doesn't make sense because there is no ON clause so you end up >>>>> with a cartesian join. Luckily, since you're looking for a specific label, >>>>> "owner", you end up with just one label and will aggregate the entire set >>>>> of files. So you end up with a result that looks reasonable, but you're >>>>> really just running two unrelated queries here: one to aggregate the total >>>>> size of live files in the table, and one to select the owner. >>>>> >>>>> I think that means that the only use case here is to expose this data >>>>> to users. But I think that this reasoning is that we need a metadata table >>>>> because we need SQL interaction and we need SQL interaction because . . . >>>>> ? >>>>> It's a nice-to-have, sure, but I'm not convinced that anyone would miss it >>>>> if we didn't expose this directly to users. >>>>> >>>>> On Fri, Sep 25, 2026 at 2:40 PM Andrei Tserakhau via dev < >>>>> [email protected]> wrote: >>>>> >>>>>> Hi Ryan, >>>>>> >>>>>> Agree here: the engine-facing consumption (cost attribution, policy >>>>>> attachment) goes through SupportsLabels, no table needed. The table is >>>>>> for >>>>>> the other consumer - SQL/people or AI Agent :). Some cases that i see >>>>>> here: >>>>>> >>>>>> 1) Exploration. "which columns are classified as X here" is a query a >>>>>> person runs, not something an engine surfaces: >>>>>> >>>>>> SELECT field_name, key, value >>>>>> FROM prod.db.orders.labels >>>>>> WHERE scope = 'field' AND key = 'classification'; >>>>>> >>>>>> 2) Composition. .labels joins with .files / .partitions / .snapshots >>>>>> in one query - e.g. attribute bytes to an owner label, i.e. cost >>>>>> attribution as a report someone runs, not a log an engine emits: >>>>>> >>>>>> SELECT l.value AS owner, SUM(f.file_size_in_bytes) AS bytes >>>>>> FROM prod.db.orders.files f, prod.db.orders.labels l >>>>>> WHERE l.scope = 'object' AND l.key = 'owner' >>>>>> GROUP BY l.value; >>>>>> >>>>>> I think the key value to have a metadata table is joinability, you >>>>>> can’t have it with a programmatic label API. >>>>>> >>>>>> So there is a place for the SQL consumer, the engine path is a >>>>>> different usecase and the co-live together. >>>>>> >>>>>> Thanks, >>>>>> Andrei >>>>>> >>>>>> >>>>>> On Fri, Sep 25, 2026 at 10:51 PM Ryan Blue <[email protected]> wrote: >>>>>> >>>>>>> > Are we OK with catalog-provided metadata tables as a separate >>>>>>> category? >>>>>>> >>>>>>> I'm okay with providing metadata through a system table like this, >>>>>>> as long as we think that people will want to access this data that way. >>>>>>> Question 3, "If no, what should the SQL surface for labels be instead?" >>>>>>> makes me think that a SQL surface is _assumed_ to be needed. >>>>>>> >>>>>>> I don't think it is necessarily the case that we need to expose >>>>>>> these for SQL users. I thought that we wanted labels to expose >>>>>>> additional context to engines for things like cost attribution logs or >>>>>>> attaching an engine's policy to a table. That doesn't require a >>>>>>> table-like >>>>>>> user surface. >>>>>>> >>>>>>> I'm fine adding a metadata table if there's a use for it, but if we >>>>>>> don't need one then it's simpler not to add and maintain it. And that >>>>>>> avoids needing to answer questions like this as well. >>>>>>> >>>>>>> >>>>>>> >>>>>>> On Fri, Sep 25, 2026 at 1:13 PM Andrei Tserakhau via dev < >>>>>>> [email protected]> wrote: >>>>>>> >>>>>>>> Hi all, >>>>>>>> >>>>>>>> Labels in the REST spec recently landed [1]. A catalog can now >>>>>>>> expose object-level and per-field labels on load-table responses. As >>>>>>>> one of >>>>>>>> the follow-ups, there is a proposal to add a .labels metadata >>>>>>>> table backed by catalog data. >>>>>>>> >>>>>>>> During discussion of the follow-ups ([3], [4]), Peter raised a good >>>>>>>> question on [3]: every metadata table today is derived from table >>>>>>>> metadata. >>>>>>>> .labels would be different because the data comes from the >>>>>>>> catalog, may vary by catalog, and may be absent if the catalog has >>>>>>>> nothing >>>>>>>> to return. >>>>>>>> >>>>>>>> I think this is less about Labels itself and more about a new >>>>>>>> *kind* of data we would expose: metadata owned by the catalog >>>>>>>> rather than by storage. Labels would be the first example, but later >>>>>>>> the >>>>>>>> same pattern could be used to expose other catalog information as >>>>>>>> metadata >>>>>>>> tables. >>>>>>>> >>>>>>>> So the broader question is: do we want metadata tables to also >>>>>>>> expose catalog-provided information? >>>>>>>> >>>>>>>> I think labels are a reasonable first case. They are structured, >>>>>>>> useful to query via SQL, and fit the same access pattern as >>>>>>>> .snapshots or .partitions. The table can stay read-only, limited >>>>>>>> to the spec-defined shape, and empty when the catalog returns nothing. >>>>>>>> >>>>>>>> The tradeoff is that this breaks the current assumption that >>>>>>>> metadata tables are deterministic projections of table metadata. It’s >>>>>>>> not >>>>>>>> explicitly written anywhere, but that assumption exists today. >>>>>>>> >>>>>>>> To summarize the questions: >>>>>>>> >>>>>>>> 1. >>>>>>>> >>>>>>>> Are we OK with catalog-provided metadata tables as a separate >>>>>>>> category? >>>>>>>> 2. >>>>>>>> >>>>>>>> If yes, should we mark them somehow so they are clearly >>>>>>>> different from spec-backed metadata tables? >>>>>>>> 3. >>>>>>>> >>>>>>>> If no, what should the SQL surface for labels be instead? >>>>>>>> >>>>>>>> Thanks, >>>>>>>> >>>>>>>> Andrei >>>>>>>> >>>>>>>> [1] REST spec labels: https://github.com/apache/iceberg/pull/15750 >>>>>>>> [2] SupportsLabels: https://github.com/apache/iceberg/pull/18046 >>>>>>>> [3] Labels metadata table: >>>>>>>> https://github.com/apache/iceberg/pull/18048 >>>>>>>> [4] Spark DESCRIBE: https://github.com/apache/iceberg/pull/18049 >>>>>>>> [5] Design: >>>>>>>> https://docs.google.com/document/d/1aj-6JlfBiMYEEVtNuh5WLMOrRQiMCcyYUGbouPM4hXI/edit?tab=t.0#heading=h.2w0kmp1v1gwv >>>>>>>> >>>>>>>
