Hi EJ, I've reviewed most of the doc (except API definitions and appendixes) and left some minor comments. Overall, the proposal looks good to me. I'll try and review the API spec PR tomorrow.
Re: namespace representation in v1: Since Polaris uses the IRC way of encoding namespaces in many other APIs, I believe it should be fine to use it in tags API v1 too. Clients will eventually query tables, which will take them to the IRC API anyway. I tend to think Polaris may need a more holistic approach to namespace encoding in _all_ REST APIs. It might require a v2 for a robust solution. A few more general comments on the proposal as a whole: * The doc does not talk about Persistence SPIs (perhaps I missed it). It would be nice to at list have a high-level desciption of the anticipated changes there. * Will lookup algorithms be delegated to Persistence implementations completely, or will Polaris have a common/shared piece of code for that, delegating lower-level simpler methods to Persistence? * It would be nice to have a dedicated section for the new Authorizer operations and their arguments. Thanks, Dmitri. On Wed, Sep 9, 2026 at 6:10 PM EJ Wang <[email protected]> wrote: > Thanks Robert, Prithvi, and Dmitri. I’ve updated the spec doc > < > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0#heading=h.jx650zq2bd2g > > > with > the contract clarifications (see below) and the follow-up discussion with > Dmitri. The corresponding PR updates are underway. > > *For encoding*, my proposal is to retain Iceberg’s namespace query > convention in v1, with explicit supported-name, encoding, and decoding > rules in Part 1 section 5.2. This retains the known separator and ingress > limitations. Since the issue also affects existing Iceberg/Polaris APIs, > I’d address a replacement codec in a separate issue/PR. Part 3 section 7.12 > records alternatives and spike results to start that discussion. > > *Version tokens* are opaque strings in both responses and update requests. > Clients return the token unchanged, and stale updates return 409. The check > must cover every supported definition-write path, while backends choose > their revision mechanism. > > *For detach-all*, the definition and assignments must disappear together > through Tag reads, or nothing changes. Physical cleanup may follow. An > implementation unable to provide that guarantee returns 501 after > authorization and before changing visible state. > > *Column assignments* use table identity and column identifier, using > Iceberg field id for Iceberg tables. Renames preserve assignments when > identity is preserved, while same-name replacements do not inherit them. V1 > remains limited to top-level Iceberg columns. > > Does this scope split work for you, particularly keeping the shared codec > redesign separate from Tag v1? I’d like to settle these contract points > here before PR1 merges. > > -ej > > On Wed, Sep 9, 2026 at 7:36 AM Dmitri Bourlatchkov <[email protected]> > wrote: > > > Hi All, > > > > (replying partially) > > > > I very much support Robert's proposal for using a well-defined and > > unambiguous format for namespaces. > > > > Many Polaris APIs fall into following the IRC approach to namespace > > representation in query parameters. Yet, that approach has multiple > issues > > , which can be seen in Iceberg dev ML / GH issues. > > > > I think Polaris should use a more robust namespace representation in its > > native APIs. > > > > Cheers, > > Dmitri. > > > > On Tue, Sep 8, 2026 at 6:41 AM Robert Stupp <[email protected]> wrote: > > > > > Hi, > > > > > > I have been thinking about what a full implementation would require. > > > > > > I like that the proposal separates definitions, assignments, and > > effective > > > reads. I have a few API-contract questions that seem worth resolving > > while > > > the contract is still separate from the implementation. > > > > > > > > > First, I think the target query parameters need a defined encoding for > > > identifier elements. > > > > > > This is not only about the unit separator. Namespace elements and > object > > > names may themselves contain characters such as &, ?, =, +, or %. > > > Those must be preserved rather than interpreted as query syntax. > > > > > > Ordinary URI query-value encoding handles those characters, but it does > > not > > > solve the separate problem of representing the boundaries between > > multipart > > > namespace elements. > > > > > > I suggest defining a small, reversible namespace-element codec, then > > > applying ordinary URI encoding to its complete output. Nessie's escaped > > > path > > > representation is a useful precedent: it has an unambiguous element > > > separator > > > and escape syntax, while avoiding control characters in the transport > > > representation. > > > > > > The contract should specify the codec, its decoding failures, and > > > conformance examples. Client libraries should expose it rather than > > > requiring > > > every client to reproduce it. The structured target used by the write > > APIs > > > would still be the clearest canonical representation; this codec would > > make > > > the GET form safe and interoperable. > > > > > > > > > Second, I think the revision token should be opaque at the API > boundary. > > > > > > The backend should be free to use a native row revision, commit ID, > ETag, > > > or > > > another conditional-write token. However, the contract should define > the > > > observable precondition: the server returns a token, and an update > > succeeds > > > only if the client supplies the token for the current tag definition. > > > Otherwise the server returns a conflict. > > > > > > That requires token matching semantics, but not an integer type, an > > initial > > > value, ordering, increment-by-one behavior, or history semantics. A > > > catalog-wide commit token would also be valid, although it could create > > > avoidable conflicts for unrelated changes. > > > > > > > > > Third, the direct reverse lookup is useful, but I would treat it as a > > > first-class, paginated relationship rather than a tag record > containing a > > > collection of targets. > > > > > > A common tag can legitimately be attached to a very large number of > > objects > > > or columns. A backend will normally need one forward access path for > > direct > > > assignments by target, and one reverse access path by tag/value, with > > > backend-specific partitioning or sharding. Effective assignments should > > > remain > > > computed from the target and its ancestors; materializing inherited > > > assignments onto descendants would have very different scaling > behavior. > > > > > > This also affects detach-all. Deleting an unbounded number of > assignment > > > records atomically is not a portable primitive for all backends. The > > > contract > > > should distinguish observable deletion semantics from physical cleanup, > > or > > > state the backend capability required for a synchronous detach-all > > > operation. > > > > > > > > > Finally, I agree with the V1 boundary of top-level Iceberg columns, > but I > > > would > > > keep the core tag model independent of Iceberg. For Iceberg, the > durable > > > column reference should be the field ID, with a column name used only > for > > > request-time resolution and display. Other table implementations could > > opt > > > in > > > later once they provide an equally stable field identity. This avoids > > > treating > > > a name-based column mapping as a general abstraction. > > > > > > > > > None of this requires tags to become authorization inputs in V1. It is > > > mainly > > > about leaving the assignment and read contract implementable by more > than > > > one > > > persistence model when those slices arrive. > > > > > > Thanks, > > > Robert > > > > > > > > > On Fri, Aug 28, 2026 at 2:19 AM EJ Wang < > [email protected]> > > > wrote: > > > > > > > Hi folks, > > > > > > > > A quick update on the Tag work. *Current status*: > > > > - PR1: API contract (https://github.com/apache/polaris/pull/5366): > > ready > > > > for review > > > > - *NEW! *PR2: Definition CRUD ( > > > https://github.com/apache/polaris/pull/5391 > > > > ): > > > > open as Draft, ready for review if PR1 LGTY > > > > - PR3: Assignment writes and storage: planned > > > > - PR4: Reads, inheritance, and reverse lookup: planned > > > > > > > > *More on PR2:* > > > > - PR2 makes Tag definitions usable through create, list, load, > update, > > > > rename, and delete. It intentionally stops before assignments, so > their > > > > persistence model remains open for the next slice. > > > > - The PR is stacked on #5366 and will be rebased once that PR merges. > > > > > > > > *Asks:* > > > > - For #5366, please call out any remaining API contract concerns. For > > > > #5391, I would especially appreciate feedback on the slice boundary > and > > > the > > > > decision to reuse the existing entity persistence model. > > > > - The design doc remains here: > > > > > > > > > > > > > > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing > > > > > > > > I’ll keep using this thread for new delivery slices, material status > > > > changes, and specific community asks. > > > > > > > > Thanks, > > > > -ej > > > > > > > > On Mon, Aug 24, 2026 at 5:04 PM EJ Wang < > > [email protected]> > > > > wrote: > > > > > > > > > Hi folks, > > > > > > > > > > Following up on this thread, I have opened a PR to land the public > > API > > > > > contract for Tags: https://github.com/apache/polaris/pull/5366 > > > > > > > > > > The PR defines Tag management, assignment and unassignment, direct > > and > > > > > inherited reads, and reverse lookup. V1 covers catalogs, > namespaces, > > > > > Iceberg and generic tables as whole objects, and top-level Iceberg > > > table > > > > > columns. Views, generic-table columns, nested fields, multi-value > > > > > assignments, and tag-based authorization are deferred. > > > > > > > > > > I plan to deliver the capability through four PRs that merge in > > order: > > > > the > > > > > API contract in this PR, Tag definition CRUD, assignment writes and > > > > > storage, then reads and reverse lookup. A separate follow-up will > add > > > > > grants on Tag resources to the management APIs. That grant surface > is > > > > > distinct from using Tags to control access to tagged objects, which > > > > remains > > > > > outside v1. > > > > > > > > > > The updated design doc is here: > > > > > > > > > > > > > > > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing > > > > > > > > > > The PR is currently Draft while we finish aligning on the public > > > > contract. > > > > > It is intended to merge as the first delivery slice, not remain as > a > > > > > design-only artifact. Please call out any remaining scope or > contract > > > > > concerns. If the list is aligned, I will mark it ready for review. > > > > > > > > > > Thanks, > > > > > -ej > > > > > > > > > > On Wed, Aug 12, 2026 at 2:05 PM EJ Wang < > > > [email protected]> > > > > > wrote: > > > > > > > > > >> Thanks Dmitri, these comments were very useful. > > > > >> > > > > >> I went through the three areas you called out and updated the > > proposal > > > > >> accordingly. > > > > >> > > > > >> On the permission/policy direction, *I agree the Tag model should > > > leave > > > > >> room for permissions or policies to consume tags later*, including > > the > > > > >> direction JB proposed. I am keeping that outside the v1 Tag > > contract, > > > > >> though. In v1, tags classify resources; they do not themselves > grant > > > or > > > > >> deny access. Polaris Policy looks like the closest existing > > foundation > > > > if > > > > >> we later want a portable tag-aware policy model, but I think that > > > > deserves > > > > >> a separate proposal rather than baking policy semantics into the > Tag > > > > >> storage model now. > > > > >> > > > > >> I also made the authorizer path more explicit. *A future OPA, > > Ranger, > > > or > > > > >> other authorizer could receive the target's complete effective > tags > > as > > > > >> resource attributes*. The authorization path would resolve those > > tags > > > > >> internally, applying target-types, inheritance, closest-wins, > > > > grandfathered > > > > >> values, and the same coherent-read guarantees as the Tag API. At > > > > minimum, > > > > >> the portable input can include the tag definition ID, current > name, > > > and > > > > >> selected value; provenance can be additional context. If Polaris > > > cannot > > > > >> resolve the complete effective state, authorization should fail > > closed > > > > >> rather than treat the resource as untagged. > > > > >> > > > > >> That also makes the persistence expectation on the read path > > clearer: > > > an > > > > >> implementation needs to resolve the target and relevant ancestors, > > > > obtain > > > > >> the applicable tag definitions and assignments, and produce one > > > coherent > > > > >> effective result. *Those observable semantics are the backend > > > contract; > > > > >> the physical lookup/indexing strategy is not.* > > > > >> > > > > >> On the Java interface suggestion, I added Java-shaped records for > > the > > > > >> durable logical model so the definition, target identity, and > > > assignment > > > > >> shapes are easier to review from JDBC and NoSQL perspectives. I > > > stopped > > > > >> short of proposing operation interfaces in pseudo-code, though. My > > > > current > > > > >> thinking is that we should first agree on the durable facts and > > > required > > > > >> behavior, then design the actual persistence SPI around the needs > of > > > the > > > > >> implementations. I did not want an illustrative interface in this > > > > design to > > > > >> accidentally become the persistence contract. > > > > >> > > > > >> So Part 2 now separates the two intentionally: > > > > >> > > > > >> *logical data + behavior/conformance requirements are specified; > > > > >> transaction, CAS, atomic batch, provider-native operations, and > the > > > > >> eventual Java SPI remain implementation/design choices.* > > > > >> > > > > >> Thanks again for the review, and definitely keep the comments > coming > > > :) > > > > >> > > > > >> I've updated the doc, please check it out the latest and the > > greatest: > > > > >> > > > > >> > > > > > > > > > > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0 > > > > >> > > > > >> -ej > > > > >> > > > > >> On Fri, Aug 7, 2026 at 3:39 PM Dmitri Bourlatchkov < > > [email protected]> > > > > >> wrote: > > > > >> > > > > >>> Hi EJ, JB, > > > > >>> > > > > >>> I left some comments on EJ's doc. I actually have a lot of > > comments > > > on > > > > >>> the > > > > >>> REST API design, I only posted some of them to start a discussion > > > > >>> without overloading the doc. > > > > >>> > > > > >>> Overall, I believe EJ's proposal should also allow permission > > > > assignments > > > > >>> on tags that JB proposed (eventually). We just need to clearly > > define > > > > the > > > > >>> persistence expectations for looking up related tags on the read > > > path. > > > > >>> > > > > >>> We should probably specify whether and how tags are exposed to > > > > >>> authorizers > > > > >>> (OPA, Ranger). I imagine people will want to use them in external > > > > policy > > > > >>> engines the moment the feature is available. > > > > >>> > > > > >>> On the persistence side, I believe it would be nice to define > > actual > > > > java > > > > >>> interfaces (perhaps in pseudo code) to allow easier review from > the > > > > NoSQL > > > > >>> persistence perspective (also commented in the doc). > > > > >>> > > > > >>> Cheers, > > > > >>> Dmitri. > > > > >>> > > > > >>> On Thu, Jul 30, 2026 at 12:39 AM Jean-Baptiste Onofré < > > > [email protected] > > > > > > > > > >>> wrote: > > > > >>> > > > > >>> > Hi EJ > > > > >>> > > > > > >>> > Thanks for starting this discussion. > > > > >>> > > > > > >>> > For the record, here's my initial proposal about tagging: > > > > >>> > > https://lists.apache.org/thread/nmqmmjfmocfllb71fcmyp9syc9gyn820 > > > > >>> > > > > > >>> > At that time, only Dmitri replied :) > > > > >>> > So, I would be happy to work with you on this, as I still have > > the > > > > PoC > > > > >>> > I created for my initial proposal. > > > > >>> > > > > > >>> > I will try to join the scheduled meeting (no guarantee). > > > > >>> > > > > > >>> > Regards > > > > >>> > JB > > > > >>> > > > > > >>> > On Fri, Jul 17, 2026 at 6:53 AM EJ Wang < > > > > >>> [email protected]> > > > > >>> > wrote: > > > > >>> > > > > > > >>> > > Hi folks, > > > > >>> > > > > > > >>> > > I have prepared a Google Doc > > > > >>> > > < > > > > >>> > > > > > >>> > > > > > > > > > > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing > > > > >>> > > > > > > >>> > > for the Polaris tag spec proposal. > > > > >>> > > > > > > >>> > > The goal is simple: add a native tag model to Polaris so > users > > > can > > > > >>> > classify > > > > >>> > > catalog objects, read those classifications back, and find > > > objects > > > > by > > > > >>> > tag. > > > > >>> > > > > > > >>> > > The proposal covers: > > > > >>> > > * tag definitions as catalog-scoped Polaris entities > > > > >>> > > * tag assignments on catalogs, namespaces, table-like > objects, > > > and > > > > >>> > columns > > > > >>> > > * allowed values on tag definitions > > > > >>> > > * direct and inherited tag reads > > > > >>> > > * direct by-tag lookup > > > > >>> > > * the durable model behind the API > > > > >>> > > * how this compares with the existing Polaris Policy API (tag > > > > design > > > > >>> > > referenced policy heavily, given their pattern similarity) > > > > >>> > > > > > > >>> > > Please take a look and leave comments in the doc. Let me know > > > WDYT! > > > > >>> > > > > > > >>> > > I would also like to discuss this in the July 23 community > > sync. > > > A > > > > >>> > separate > > > > >>> > > dedicated review meeting will be scheduled separately, likely > > > > within > > > > >>> the > > > > >>> > > next two weeks. > > > > >>> > > > > > > >>> > > Thanks, > > > > >>> > > -ej > > > > >>> > > > > > >>> > > > > >> > > > > > > > > > >
