Hi all, Reviving this with a narrower question after the same issue came up again in iceberg-go [1].
The java labels metadata-table PR [0] was closed because metadata tables in Java/Spark are expected to be deterministic projections of table metadata. I think that still makes sense for the engine SQL / Spark surface. The non-Java clients are a bit different, though. In PyIceberg, iceberg-rust, and iceberg-go, inspect is a library read API, not a SQL metadata table, and there is no convinient alternative. My suggestion is to keep the Java/Spark decision as-is, but let each client expose catalog-provided data through its inspect API if useful. For labels, that means documenting it as catalog-provided, captured at load time, and empty when the catalog returns none. If that split sounds reasonable, I'll proceed with the iceberg-go PR [1] and mirror it in rust and python. Any objections to treating the client inspect API separately from the engine metadata-table contract? Thanks, Andrei [0] https://github.com/apache/iceberg/pull/18048 (Java core labels metadata table, closed) [1] https://github.com/apache/iceberg-go/pull/2101 (iceberg-go labels inspect table) On Tue, Sep 29, 2026 at 12:19 AM Andrei Tserakhau < [email protected]> wrote: > Fair - I think this brings the focus back to the main question: why do we > need > SQL at all. > > My answer is that there's a class of consumers whose only interface is SQL > - > SQL-native discovery / governance tooling, and most notably LLM agents that > explore a warehouse through SQL. For them a Java capability > (SupportsLabels) is > not a thing; they can only reason over what they can query. > > What can cover that is `DESCRIBE` - #18049 surfaces both object and field > labels > there (labels.object.*, labels.field.<id>.*). So if there's no urge to join > stuff, I see the point - we can keep it simple and not do the metadata > table. > > We can revisit this topic later, once we have more data points and use > cases, > but for now I agree - it's not needed. > > Best, > Andrei > > On Mon, Sep 28, 2026 at 11:30 PM Ryan Blue <[email protected]> wrote: > >> I don't agree with the composition argument. The example query you >> provided doesn't make sense because there is no ON clause so you end up >> with a cartesian join. Luckily, since you're looking for a specific label, >> "owner", you end up with just one label and will aggregate the entire set >> of files. So you end up with a result that looks reasonable, but you're >> really just running two unrelated queries here: one to aggregate the total >> size of live files in the table, and one to select the owner. >> >> I think that means that the only use case here is to expose this data to >> users. But I think that this reasoning is that we need a metadata table >> because we need SQL interaction and we need SQL interaction because . . . ? >> It's a nice-to-have, sure, but I'm not convinced that anyone would miss it >> if we didn't expose this directly to users. >> >> On Fri, Sep 25, 2026 at 2:40 PM Andrei Tserakhau via dev < >> [email protected]> wrote: >> >>> Hi Ryan, >>> >>> Agree here: the engine-facing consumption (cost attribution, policy >>> attachment) goes through SupportsLabels, no table needed. The table is for >>> the other consumer - SQL/people or AI Agent :). Some cases that i see here: >>> >>> 1) Exploration. "which columns are classified as X here" is a query a >>> person runs, not something an engine surfaces: >>> >>> SELECT field_name, key, value >>> FROM prod.db.orders.labels >>> WHERE scope = 'field' AND key = 'classification'; >>> >>> 2) Composition. .labels joins with .files / .partitions / .snapshots in >>> one query - e.g. attribute bytes to an owner label, i.e. cost attribution >>> as a report someone runs, not a log an engine emits: >>> >>> SELECT l.value AS owner, SUM(f.file_size_in_bytes) AS bytes >>> FROM prod.db.orders.files f, prod.db.orders.labels l >>> WHERE l.scope = 'object' AND l.key = 'owner' >>> GROUP BY l.value; >>> >>> I think the key value to have a metadata table is joinability, you can’t >>> have it with a programmatic label API. >>> >>> So there is a place for the SQL consumer, the engine path is a different >>> usecase and the co-live together. >>> >>> Thanks, >>> Andrei >>> >>> >>> On Fri, Sep 25, 2026 at 10:51 PM Ryan Blue <[email protected]> wrote: >>> >>>> > Are we OK with catalog-provided metadata tables as a separate >>>> category? >>>> >>>> I'm okay with providing metadata through a system table like this, as >>>> long as we think that people will want to access this data that way. >>>> Question 3, "If no, what should the SQL surface for labels be instead?" >>>> makes me think that a SQL surface is _assumed_ to be needed. >>>> >>>> I don't think it is necessarily the case that we need to expose these >>>> for SQL users. I thought that we wanted labels to expose additional context >>>> to engines for things like cost attribution logs or attaching an engine's >>>> policy to a table. That doesn't require a table-like user surface. >>>> >>>> I'm fine adding a metadata table if there's a use for it, but if we >>>> don't need one then it's simpler not to add and maintain it. And that >>>> avoids needing to answer questions like this as well. >>>> >>>> >>>> >>>> On Fri, Sep 25, 2026 at 1:13 PM Andrei Tserakhau via dev < >>>> [email protected]> wrote: >>>> >>>>> Hi all, >>>>> >>>>> Labels in the REST spec recently landed [1]. A catalog can now expose >>>>> object-level and per-field labels on load-table responses. As one of the >>>>> follow-ups, there is a proposal to add a .labels metadata table >>>>> backed by catalog data. >>>>> >>>>> During discussion of the follow-ups ([3], [4]), Peter raised a good >>>>> question on [3]: every metadata table today is derived from table >>>>> metadata. >>>>> .labels would be different because the data comes from the catalog, >>>>> may vary by catalog, and may be absent if the catalog has nothing to >>>>> return. >>>>> >>>>> I think this is less about Labels itself and more about a new *kind* >>>>> of data we would expose: metadata owned by the catalog rather than by >>>>> storage. Labels would be the first example, but later the same pattern >>>>> could be used to expose other catalog information as metadata tables. >>>>> >>>>> So the broader question is: do we want metadata tables to also expose >>>>> catalog-provided information? >>>>> >>>>> I think labels are a reasonable first case. They are structured, >>>>> useful to query via SQL, and fit the same access pattern as .snapshots >>>>> or .partitions. The table can stay read-only, limited to the >>>>> spec-defined shape, and empty when the catalog returns nothing. >>>>> >>>>> The tradeoff is that this breaks the current assumption that metadata >>>>> tables are deterministic projections of table metadata. It’s not >>>>> explicitly >>>>> written anywhere, but that assumption exists today. >>>>> >>>>> To summarize the questions: >>>>> >>>>> 1. >>>>> >>>>> Are we OK with catalog-provided metadata tables as a separate >>>>> category? >>>>> 2. >>>>> >>>>> If yes, should we mark them somehow so they are clearly different >>>>> from spec-backed metadata tables? >>>>> 3. >>>>> >>>>> If no, what should the SQL surface for labels be instead? >>>>> >>>>> Thanks, >>>>> >>>>> Andrei >>>>> >>>>> [1] REST spec labels: https://github.com/apache/iceberg/pull/15750 >>>>> [2] SupportsLabels: https://github.com/apache/iceberg/pull/18046 >>>>> [3] Labels metadata table: >>>>> https://github.com/apache/iceberg/pull/18048 >>>>> [4] Spark DESCRIBE: https://github.com/apache/iceberg/pull/18049 >>>>> [5] Design: >>>>> https://docs.google.com/document/d/1aj-6JlfBiMYEEVtNuh5WLMOrRQiMCcyYUGbouPM4hXI/edit?tab=t.0#heading=h.2w0kmp1v1gwv >>>>> >>>>
