Hi folks, A quick update on the Tag work. PR1-3 and the design doc have been updated following the discussion.
*Current status*: PR1: API contract: ready for review. PR2: Definition CRUD: open as Draft, ready for review if PR1 LGTY. *NEW! *PR3: Assignment writes and storage: open as Draft, ready for review if PR2 LGTY. PR4: Reads, inheritance, and reverse lookup: planned. *More on PR3*: PR3 adds assign/unassign for catalogs, namespaces, tables, and top-level Iceberg columns, plus atomic detach-all. It includes JDBC and in-memory support. Reads remain in PR4. The PR is stacked on PR2. Its description links the assignment-only commit for focused review. *Asks*: For PR3, I’d especially appreciate feedback on the persistence interfaces, their impact on external implementations, and the detach-all guarantee. The updated design doc also includes the reverse-lookup flow and a discussion of the deferred shared encoding work. Thanks, -ej On Wed, Sep 9, 2026 at 3:09 PM EJ Wang <[email protected]> wrote: > Thanks Robert, Prithvi, and Dmitri. I’ve updated the spec doc > <https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0#heading=h.jx650zq2bd2g> > with > the contract clarifications (see below) and the follow-up discussion with > Dmitri. The corresponding PR updates are underway. > > *For encoding*, my proposal is to retain Iceberg’s namespace query > convention in v1, with explicit supported-name, encoding, and decoding > rules in Part 1 section 5.2. This retains the known separator and ingress > limitations. Since the issue also affects existing Iceberg/Polaris APIs, > I’d address a replacement codec in a separate issue/PR. Part 3 section 7.12 > records alternatives and spike results to start that discussion. > > *Version tokens* are opaque strings in both responses and update > requests. Clients return the token unchanged, and stale updates return 409. > The check must cover every supported definition-write path, while backends > choose their revision mechanism. > > *For detach-all*, the definition and assignments must disappear together > through Tag reads, or nothing changes. Physical cleanup may follow. An > implementation unable to provide that guarantee returns 501 after > authorization and before changing visible state. > > *Column assignments* use table identity and column identifier, using > Iceberg field id for Iceberg tables. Renames preserve assignments when > identity is preserved, while same-name replacements do not inherit them. V1 > remains limited to top-level Iceberg columns. > > Does this scope split work for you, particularly keeping the shared codec > redesign separate from Tag v1? I’d like to settle these contract points > here before PR1 merges. > > -ej > > On Wed, Sep 9, 2026 at 7:36 AM Dmitri Bourlatchkov <[email protected]> > wrote: > >> Hi All, >> >> (replying partially) >> >> I very much support Robert's proposal for using a well-defined and >> unambiguous format for namespaces. >> >> Many Polaris APIs fall into following the IRC approach to namespace >> representation in query parameters. Yet, that approach has multiple issues >> , which can be seen in Iceberg dev ML / GH issues. >> >> I think Polaris should use a more robust namespace representation in its >> native APIs. >> >> Cheers, >> Dmitri. >> >> On Tue, Sep 8, 2026 at 6:41 AM Robert Stupp <[email protected]> wrote: >> >> > Hi, >> > >> > I have been thinking about what a full implementation would require. >> > >> > I like that the proposal separates definitions, assignments, and >> effective >> > reads. I have a few API-contract questions that seem worth resolving >> while >> > the contract is still separate from the implementation. >> > >> > >> > First, I think the target query parameters need a defined encoding for >> > identifier elements. >> > >> > This is not only about the unit separator. Namespace elements and object >> > names may themselves contain characters such as &, ?, =, +, or %. >> > Those must be preserved rather than interpreted as query syntax. >> > >> > Ordinary URI query-value encoding handles those characters, but it does >> not >> > solve the separate problem of representing the boundaries between >> multipart >> > namespace elements. >> > >> > I suggest defining a small, reversible namespace-element codec, then >> > applying ordinary URI encoding to its complete output. Nessie's escaped >> > path >> > representation is a useful precedent: it has an unambiguous element >> > separator >> > and escape syntax, while avoiding control characters in the transport >> > representation. >> > >> > The contract should specify the codec, its decoding failures, and >> > conformance examples. Client libraries should expose it rather than >> > requiring >> > every client to reproduce it. The structured target used by the write >> APIs >> > would still be the clearest canonical representation; this codec would >> make >> > the GET form safe and interoperable. >> > >> > >> > Second, I think the revision token should be opaque at the API boundary. >> > >> > The backend should be free to use a native row revision, commit ID, >> ETag, >> > or >> > another conditional-write token. However, the contract should define the >> > observable precondition: the server returns a token, and an update >> succeeds >> > only if the client supplies the token for the current tag definition. >> > Otherwise the server returns a conflict. >> > >> > That requires token matching semantics, but not an integer type, an >> initial >> > value, ordering, increment-by-one behavior, or history semantics. A >> > catalog-wide commit token would also be valid, although it could create >> > avoidable conflicts for unrelated changes. >> > >> > >> > Third, the direct reverse lookup is useful, but I would treat it as a >> > first-class, paginated relationship rather than a tag record containing >> a >> > collection of targets. >> > >> > A common tag can legitimately be attached to a very large number of >> objects >> > or columns. A backend will normally need one forward access path for >> direct >> > assignments by target, and one reverse access path by tag/value, with >> > backend-specific partitioning or sharding. Effective assignments should >> > remain >> > computed from the target and its ancestors; materializing inherited >> > assignments onto descendants would have very different scaling behavior. >> > >> > This also affects detach-all. Deleting an unbounded number of assignment >> > records atomically is not a portable primitive for all backends. The >> > contract >> > should distinguish observable deletion semantics from physical cleanup, >> or >> > state the backend capability required for a synchronous detach-all >> > operation. >> > >> > >> > Finally, I agree with the V1 boundary of top-level Iceberg columns, but >> I >> > would >> > keep the core tag model independent of Iceberg. For Iceberg, the durable >> > column reference should be the field ID, with a column name used only >> for >> > request-time resolution and display. Other table implementations could >> opt >> > in >> > later once they provide an equally stable field identity. This avoids >> > treating >> > a name-based column mapping as a general abstraction. >> > >> > >> > None of this requires tags to become authorization inputs in V1. It is >> > mainly >> > about leaving the assignment and read contract implementable by more >> than >> > one >> > persistence model when those slices arrive. >> > >> > Thanks, >> > Robert >> > >> > >> > On Fri, Aug 28, 2026 at 2:19 AM EJ Wang <[email protected] >> > >> > wrote: >> > >> > > Hi folks, >> > > >> > > A quick update on the Tag work. *Current status*: >> > > - PR1: API contract (https://github.com/apache/polaris/pull/5366): >> ready >> > > for review >> > > - *NEW! *PR2: Definition CRUD ( >> > https://github.com/apache/polaris/pull/5391 >> > > ): >> > > open as Draft, ready for review if PR1 LGTY >> > > - PR3: Assignment writes and storage: planned >> > > - PR4: Reads, inheritance, and reverse lookup: planned >> > > >> > > *More on PR2:* >> > > - PR2 makes Tag definitions usable through create, list, load, update, >> > > rename, and delete. It intentionally stops before assignments, so >> their >> > > persistence model remains open for the next slice. >> > > - The PR is stacked on #5366 and will be rebased once that PR merges. >> > > >> > > *Asks:* >> > > - For #5366, please call out any remaining API contract concerns. For >> > > #5391, I would especially appreciate feedback on the slice boundary >> and >> > the >> > > decision to reuse the existing entity persistence model. >> > > - The design doc remains here: >> > > >> > > >> > >> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing >> > > >> > > I’ll keep using this thread for new delivery slices, material status >> > > changes, and specific community asks. >> > > >> > > Thanks, >> > > -ej >> > > >> > > On Mon, Aug 24, 2026 at 5:04 PM EJ Wang < >> [email protected]> >> > > wrote: >> > > >> > > > Hi folks, >> > > > >> > > > Following up on this thread, I have opened a PR to land the public >> API >> > > > contract for Tags: https://github.com/apache/polaris/pull/5366 >> > > > >> > > > The PR defines Tag management, assignment and unassignment, direct >> and >> > > > inherited reads, and reverse lookup. V1 covers catalogs, namespaces, >> > > > Iceberg and generic tables as whole objects, and top-level Iceberg >> > table >> > > > columns. Views, generic-table columns, nested fields, multi-value >> > > > assignments, and tag-based authorization are deferred. >> > > > >> > > > I plan to deliver the capability through four PRs that merge in >> order: >> > > the >> > > > API contract in this PR, Tag definition CRUD, assignment writes and >> > > > storage, then reads and reverse lookup. A separate follow-up will >> add >> > > > grants on Tag resources to the management APIs. That grant surface >> is >> > > > distinct from using Tags to control access to tagged objects, which >> > > remains >> > > > outside v1. >> > > > >> > > > The updated design doc is here: >> > > > >> > > >> > >> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing >> > > > >> > > > The PR is currently Draft while we finish aligning on the public >> > > contract. >> > > > It is intended to merge as the first delivery slice, not remain as a >> > > > design-only artifact. Please call out any remaining scope or >> contract >> > > > concerns. If the list is aligned, I will mark it ready for review. >> > > > >> > > > Thanks, >> > > > -ej >> > > > >> > > > On Wed, Aug 12, 2026 at 2:05 PM EJ Wang < >> > [email protected]> >> > > > wrote: >> > > > >> > > >> Thanks Dmitri, these comments were very useful. >> > > >> >> > > >> I went through the three areas you called out and updated the >> proposal >> > > >> accordingly. >> > > >> >> > > >> On the permission/policy direction, *I agree the Tag model should >> > leave >> > > >> room for permissions or policies to consume tags later*, including >> the >> > > >> direction JB proposed. I am keeping that outside the v1 Tag >> contract, >> > > >> though. In v1, tags classify resources; they do not themselves >> grant >> > or >> > > >> deny access. Polaris Policy looks like the closest existing >> foundation >> > > if >> > > >> we later want a portable tag-aware policy model, but I think that >> > > deserves >> > > >> a separate proposal rather than baking policy semantics into the >> Tag >> > > >> storage model now. >> > > >> >> > > >> I also made the authorizer path more explicit. *A future OPA, >> Ranger, >> > or >> > > >> other authorizer could receive the target's complete effective >> tags as >> > > >> resource attributes*. The authorization path would resolve those >> tags >> > > >> internally, applying target-types, inheritance, closest-wins, >> > > grandfathered >> > > >> values, and the same coherent-read guarantees as the Tag API. At >> > > minimum, >> > > >> the portable input can include the tag definition ID, current name, >> > and >> > > >> selected value; provenance can be additional context. If Polaris >> > cannot >> > > >> resolve the complete effective state, authorization should fail >> closed >> > > >> rather than treat the resource as untagged. >> > > >> >> > > >> That also makes the persistence expectation on the read path >> clearer: >> > an >> > > >> implementation needs to resolve the target and relevant ancestors, >> > > obtain >> > > >> the applicable tag definitions and assignments, and produce one >> > coherent >> > > >> effective result. *Those observable semantics are the backend >> > contract; >> > > >> the physical lookup/indexing strategy is not.* >> > > >> >> > > >> On the Java interface suggestion, I added Java-shaped records for >> the >> > > >> durable logical model so the definition, target identity, and >> > assignment >> > > >> shapes are easier to review from JDBC and NoSQL perspectives. I >> > stopped >> > > >> short of proposing operation interfaces in pseudo-code, though. My >> > > current >> > > >> thinking is that we should first agree on the durable facts and >> > required >> > > >> behavior, then design the actual persistence SPI around the needs >> of >> > the >> > > >> implementations. I did not want an illustrative interface in this >> > > design to >> > > >> accidentally become the persistence contract. >> > > >> >> > > >> So Part 2 now separates the two intentionally: >> > > >> >> > > >> *logical data + behavior/conformance requirements are specified; >> > > >> transaction, CAS, atomic batch, provider-native operations, and the >> > > >> eventual Java SPI remain implementation/design choices.* >> > > >> >> > > >> Thanks again for the review, and definitely keep the comments >> coming >> > :) >> > > >> >> > > >> I've updated the doc, please check it out the latest and the >> greatest: >> > > >> >> > > >> >> > > >> > >> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0 >> > > >> >> > > >> -ej >> > > >> >> > > >> On Fri, Aug 7, 2026 at 3:39 PM Dmitri Bourlatchkov < >> [email protected]> >> > > >> wrote: >> > > >> >> > > >>> Hi EJ, JB, >> > > >>> >> > > >>> I left some comments on EJ's doc. I actually have a lot of >> comments >> > on >> > > >>> the >> > > >>> REST API design, I only posted some of them to start a discussion >> > > >>> without overloading the doc. >> > > >>> >> > > >>> Overall, I believe EJ's proposal should also allow permission >> > > assignments >> > > >>> on tags that JB proposed (eventually). We just need to clearly >> define >> > > the >> > > >>> persistence expectations for looking up related tags on the read >> > path. >> > > >>> >> > > >>> We should probably specify whether and how tags are exposed to >> > > >>> authorizers >> > > >>> (OPA, Ranger). I imagine people will want to use them in external >> > > policy >> > > >>> engines the moment the feature is available. >> > > >>> >> > > >>> On the persistence side, I believe it would be nice to define >> actual >> > > java >> > > >>> interfaces (perhaps in pseudo code) to allow easier review from >> the >> > > NoSQL >> > > >>> persistence perspective (also commented in the doc). >> > > >>> >> > > >>> Cheers, >> > > >>> Dmitri. >> > > >>> >> > > >>> On Thu, Jul 30, 2026 at 12:39 AM Jean-Baptiste Onofré < >> > [email protected] >> > > > >> > > >>> wrote: >> > > >>> >> > > >>> > Hi EJ >> > > >>> > >> > > >>> > Thanks for starting this discussion. >> > > >>> > >> > > >>> > For the record, here's my initial proposal about tagging: >> > > >>> > >> https://lists.apache.org/thread/nmqmmjfmocfllb71fcmyp9syc9gyn820 >> > > >>> > >> > > >>> > At that time, only Dmitri replied :) >> > > >>> > So, I would be happy to work with you on this, as I still have >> the >> > > PoC >> > > >>> > I created for my initial proposal. >> > > >>> > >> > > >>> > I will try to join the scheduled meeting (no guarantee). >> > > >>> > >> > > >>> > Regards >> > > >>> > JB >> > > >>> > >> > > >>> > On Fri, Jul 17, 2026 at 6:53 AM EJ Wang < >> > > >>> [email protected]> >> > > >>> > wrote: >> > > >>> > > >> > > >>> > > Hi folks, >> > > >>> > > >> > > >>> > > I have prepared a Google Doc >> > > >>> > > < >> > > >>> > >> > > >>> >> > > >> > >> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing >> > > >>> > > >> > > >>> > > for the Polaris tag spec proposal. >> > > >>> > > >> > > >>> > > The goal is simple: add a native tag model to Polaris so users >> > can >> > > >>> > classify >> > > >>> > > catalog objects, read those classifications back, and find >> > objects >> > > by >> > > >>> > tag. >> > > >>> > > >> > > >>> > > The proposal covers: >> > > >>> > > * tag definitions as catalog-scoped Polaris entities >> > > >>> > > * tag assignments on catalogs, namespaces, table-like objects, >> > and >> > > >>> > columns >> > > >>> > > * allowed values on tag definitions >> > > >>> > > * direct and inherited tag reads >> > > >>> > > * direct by-tag lookup >> > > >>> > > * the durable model behind the API >> > > >>> > > * how this compares with the existing Polaris Policy API (tag >> > > design >> > > >>> > > referenced policy heavily, given their pattern similarity) >> > > >>> > > >> > > >>> > > Please take a look and leave comments in the doc. Let me know >> > WDYT! >> > > >>> > > >> > > >>> > > I would also like to discuss this in the July 23 community >> sync. >> > A >> > > >>> > separate >> > > >>> > > dedicated review meeting will be scheduled separately, likely >> > > within >> > > >>> the >> > > >>> > > next two weeks. >> > > >>> > > >> > > >>> > > Thanks, >> > > >>> > > -ej >> > > >>> > >> > > >>> >> > > >> >> > > >> > >> >
