Hi folks,

A quick update on the Tag work. PR1-3 and the design doc have been updated
following the discussion.

*Current status*:
PR1: API contract: ready for review.
PR2: Definition CRUD: open as Draft, ready for review if PR1 LGTY.
*NEW! *PR3: Assignment writes and storage: open as Draft, ready for review
if PR2 LGTY.
PR4: Reads, inheritance, and reverse lookup: planned.

*More on PR3*:
PR3 adds assign/unassign for catalogs, namespaces, tables, and top-level
Iceberg columns, plus atomic detach-all. It includes JDBC and in-memory
support. Reads remain in PR4.
The PR is stacked on PR2. Its description links the assignment-only commit
for focused review.

*Asks*:
For PR3, I’d especially appreciate feedback on the persistence interfaces,
their impact on external implementations, and the detach-all guarantee.
The updated design doc also includes the reverse-lookup flow and a
discussion of the deferred shared encoding work.

Thanks,
-ej

On Wed, Sep 9, 2026 at 3:09 PM EJ Wang <[email protected]>
wrote:

> Thanks Robert, Prithvi, and Dmitri. I’ve updated the spec doc
> <https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0#heading=h.jx650zq2bd2g>
>  with
> the contract clarifications (see below) and the follow-up discussion with
> Dmitri. The corresponding PR updates are underway.
>
> *For encoding*, my proposal is to retain Iceberg’s namespace query
> convention in v1, with explicit supported-name, encoding, and decoding
> rules in Part 1 section 5.2. This retains the known separator and ingress
> limitations. Since the issue also affects existing Iceberg/Polaris APIs,
> I’d address a replacement codec in a separate issue/PR. Part 3 section 7.12
> records alternatives and spike results to start that discussion.
>
> *Version tokens* are opaque strings in both responses and update
> requests. Clients return the token unchanged, and stale updates return 409.
> The check must cover every supported definition-write path, while backends
> choose their revision mechanism.
>
> *For detach-all*, the definition and assignments must disappear together
> through Tag reads, or nothing changes. Physical cleanup may follow. An
> implementation unable to provide that guarantee returns 501 after
> authorization and before changing visible state.
>
> *Column assignments* use table identity and column identifier, using
> Iceberg field id for Iceberg tables. Renames preserve assignments when
> identity is preserved, while same-name replacements do not inherit them. V1
> remains limited to top-level Iceberg columns.
>
> Does this scope split work for you, particularly keeping the shared codec
> redesign separate from Tag v1? I’d like to settle these contract points
> here before PR1 merges.
>
> -ej
>
> On Wed, Sep 9, 2026 at 7:36 AM Dmitri Bourlatchkov <[email protected]>
> wrote:
>
>> Hi All,
>>
>> (replying partially)
>>
>> I very much support Robert's proposal for using a well-defined and
>> unambiguous format for namespaces.
>>
>> Many Polaris APIs fall into following the IRC approach to namespace
>> representation in query parameters. Yet, that approach has multiple issues
>> , which can be seen in Iceberg dev ML / GH issues.
>>
>> I think Polaris should use a more robust namespace representation in its
>> native APIs.
>>
>> Cheers,
>> Dmitri.
>>
>> On Tue, Sep 8, 2026 at 6:41 AM Robert Stupp <[email protected]> wrote:
>>
>> > Hi,
>> >
>> > I have been thinking about what a full implementation would require.
>> >
>> > I like that the proposal separates definitions, assignments, and
>> effective
>> > reads. I have a few API-contract questions that seem worth resolving
>> while
>> > the contract is still separate from the implementation.
>> >
>> >
>> > First, I think the target query parameters need a defined encoding for
>> > identifier elements.
>> >
>> > This is not only about the unit separator. Namespace elements and object
>> > names may themselves contain characters such as &, ?, =, +, or %.
>> > Those must be preserved rather than interpreted as query syntax.
>> >
>> > Ordinary URI query-value encoding handles those characters, but it does
>> not
>> > solve the separate problem of representing the boundaries between
>> multipart
>> > namespace elements.
>> >
>> > I suggest defining a small, reversible namespace-element codec, then
>> > applying ordinary URI encoding to its complete output. Nessie's escaped
>> > path
>> > representation is a useful precedent: it has an unambiguous element
>> > separator
>> > and escape syntax, while avoiding control characters in the transport
>> > representation.
>> >
>> > The contract should specify the codec, its decoding failures, and
>> > conformance examples. Client libraries should expose it rather than
>> > requiring
>> > every client to reproduce it. The structured target used by the write
>> APIs
>> > would still be the clearest canonical representation; this codec would
>> make
>> > the GET form safe and interoperable.
>> >
>> >
>> > Second, I think the revision token should be opaque at the API boundary.
>> >
>> > The backend should be free to use a native row revision, commit ID,
>> ETag,
>> > or
>> > another conditional-write token. However, the contract should define the
>> > observable precondition: the server returns a token, and an update
>> succeeds
>> > only if the client supplies the token for the current tag definition.
>> > Otherwise the server returns a conflict.
>> >
>> > That requires token matching semantics, but not an integer type, an
>> initial
>> > value, ordering, increment-by-one behavior, or history semantics. A
>> > catalog-wide commit token would also be valid, although it could create
>> > avoidable conflicts for unrelated changes.
>> >
>> >
>> > Third, the direct reverse lookup is useful, but I would treat it as a
>> > first-class, paginated relationship rather than a tag record containing
>> a
>> > collection of targets.
>> >
>> > A common tag can legitimately be attached to a very large number of
>> objects
>> > or columns. A backend will normally need one forward access path for
>> direct
>> > assignments by target, and one reverse access path by tag/value, with
>> > backend-specific partitioning or sharding. Effective assignments should
>> > remain
>> > computed from the target and its ancestors; materializing inherited
>> > assignments onto descendants would have very different scaling behavior.
>> >
>> > This also affects detach-all. Deleting an unbounded number of assignment
>> > records atomically is not a portable primitive for all backends. The
>> > contract
>> > should distinguish observable deletion semantics from physical cleanup,
>> or
>> > state the backend capability required for a synchronous detach-all
>> > operation.
>> >
>> >
>> > Finally, I agree with the V1 boundary of top-level Iceberg columns, but
>> I
>> > would
>> > keep the core tag model independent of Iceberg. For Iceberg, the durable
>> > column reference should be the field ID, with a column name used only
>> for
>> > request-time resolution and display. Other table implementations could
>> opt
>> > in
>> > later once they provide an equally stable field identity. This avoids
>> > treating
>> > a name-based column mapping as a general abstraction.
>> >
>> >
>> > None of this requires tags to become authorization inputs in V1. It is
>> > mainly
>> > about leaving the assignment and read contract implementable by more
>> than
>> > one
>> > persistence model when those slices arrive.
>> >
>> > Thanks,
>> > Robert
>> >
>> >
>> > On Fri, Aug 28, 2026 at 2:19 AM EJ Wang <[email protected]
>> >
>> > wrote:
>> >
>> > > Hi folks,
>> > >
>> > > A quick update on the Tag work. *Current status*:
>> > > - PR1: API contract (https://github.com/apache/polaris/pull/5366):
>> ready
>> > > for review
>> > > - *NEW! *PR2: Definition CRUD (
>> > https://github.com/apache/polaris/pull/5391
>> > > ):
>> > > open as Draft, ready for review if PR1 LGTY
>> > > - PR3: Assignment writes and storage: planned
>> > > - PR4: Reads, inheritance, and reverse lookup: planned
>> > >
>> > > *More on PR2:*
>> > > - PR2 makes Tag definitions usable through create, list, load, update,
>> > > rename, and delete. It intentionally stops before assignments, so
>> their
>> > > persistence model remains open for the next slice.
>> > > - The PR is stacked on #5366 and will be rebased once that PR merges.
>> > >
>> > > *Asks:*
>> > > - For #5366, please call out any remaining API contract concerns. For
>> > > #5391, I would especially appreciate feedback on the slice boundary
>> and
>> > the
>> > > decision to reuse the existing entity persistence model.
>> > > - The design doc remains here:
>> > >
>> > >
>> >
>> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
>> > >
>> > > I’ll keep using this thread for new delivery slices, material status
>> > > changes, and specific community asks.
>> > >
>> > > Thanks,
>> > > -ej
>> > >
>> > > On Mon, Aug 24, 2026 at 5:04 PM EJ Wang <
>> [email protected]>
>> > > wrote:
>> > >
>> > > > Hi folks,
>> > > >
>> > > > Following up on this thread, I have opened a PR to land the public
>> API
>> > > > contract for Tags: https://github.com/apache/polaris/pull/5366
>> > > >
>> > > > The PR defines Tag management, assignment and unassignment, direct
>> and
>> > > > inherited reads, and reverse lookup. V1 covers catalogs, namespaces,
>> > > > Iceberg and generic tables as whole objects, and top-level Iceberg
>> > table
>> > > > columns. Views, generic-table columns, nested fields, multi-value
>> > > > assignments, and tag-based authorization are deferred.
>> > > >
>> > > > I plan to deliver the capability through four PRs that merge in
>> order:
>> > > the
>> > > > API contract in this PR, Tag definition CRUD, assignment writes and
>> > > > storage, then reads and reverse lookup. A separate follow-up will
>> add
>> > > > grants on Tag resources to the management APIs. That grant surface
>> is
>> > > > distinct from using Tags to control access to tagged objects, which
>> > > remains
>> > > > outside v1.
>> > > >
>> > > > The updated design doc is here:
>> > > >
>> > >
>> >
>> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
>> > > >
>> > > > The PR is currently Draft while we finish aligning on the public
>> > > contract.
>> > > > It is intended to merge as the first delivery slice, not remain as a
>> > > > design-only artifact. Please call out any remaining scope or
>> contract
>> > > > concerns. If the list is aligned, I will mark it ready for review.
>> > > >
>> > > > Thanks,
>> > > > -ej
>> > > >
>> > > > On Wed, Aug 12, 2026 at 2:05 PM EJ Wang <
>> > [email protected]>
>> > > > wrote:
>> > > >
>> > > >> Thanks Dmitri, these comments were very useful.
>> > > >>
>> > > >> I went through the three areas you called out and updated the
>> proposal
>> > > >> accordingly.
>> > > >>
>> > > >> On the permission/policy direction, *I agree the Tag model should
>> > leave
>> > > >> room for permissions or policies to consume tags later*, including
>> the
>> > > >> direction JB proposed. I am keeping that outside the v1 Tag
>> contract,
>> > > >> though. In v1, tags classify resources; they do not themselves
>> grant
>> > or
>> > > >> deny access. Polaris Policy looks like the closest existing
>> foundation
>> > > if
>> > > >> we later want a portable tag-aware policy model, but I think that
>> > > deserves
>> > > >> a separate proposal rather than baking policy semantics into the
>> Tag
>> > > >> storage model now.
>> > > >>
>> > > >> I also made the authorizer path more explicit. *A future OPA,
>> Ranger,
>> > or
>> > > >> other authorizer could receive the target's complete effective
>> tags as
>> > > >> resource attributes*. The authorization path would resolve those
>> tags
>> > > >> internally, applying target-types, inheritance, closest-wins,
>> > > grandfathered
>> > > >> values, and the same coherent-read guarantees as the Tag API. At
>> > > minimum,
>> > > >> the portable input can include the tag definition ID, current name,
>> > and
>> > > >> selected value; provenance can be additional context. If Polaris
>> > cannot
>> > > >> resolve the complete effective state, authorization should fail
>> closed
>> > > >> rather than treat the resource as untagged.
>> > > >>
>> > > >> That also makes the persistence expectation on the read path
>> clearer:
>> > an
>> > > >> implementation needs to resolve the target and relevant ancestors,
>> > > obtain
>> > > >> the applicable tag definitions and assignments, and produce one
>> > coherent
>> > > >> effective result. *Those observable semantics are the backend
>> > contract;
>> > > >> the physical lookup/indexing strategy is not.*
>> > > >>
>> > > >> On the Java interface suggestion, I added Java-shaped records for
>> the
>> > > >> durable logical model so the definition, target identity, and
>> > assignment
>> > > >> shapes are easier to review from JDBC and NoSQL perspectives. I
>> > stopped
>> > > >> short of proposing operation interfaces in pseudo-code, though. My
>> > > current
>> > > >> thinking is that we should first agree on the durable facts and
>> > required
>> > > >> behavior, then design the actual persistence SPI around the needs
>> of
>> > the
>> > > >> implementations. I did not want an illustrative interface in this
>> > > design to
>> > > >> accidentally become the persistence contract.
>> > > >>
>> > > >> So Part 2 now separates the two intentionally:
>> > > >>
>> > > >> *logical data + behavior/conformance requirements are specified;
>> > > >> transaction, CAS, atomic batch, provider-native operations, and the
>> > > >> eventual Java SPI remain implementation/design choices.*
>> > > >>
>> > > >> Thanks again for the review, and definitely keep the comments
>> coming
>> > :)
>> > > >>
>> > > >> I've updated the doc, please check it out the latest and the
>> greatest:
>> > > >>
>> > > >>
>> > >
>> >
>> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0
>> > > >>
>> > > >> -ej
>> > > >>
>> > > >> On Fri, Aug 7, 2026 at 3:39 PM Dmitri Bourlatchkov <
>> [email protected]>
>> > > >> wrote:
>> > > >>
>> > > >>> Hi EJ, JB,
>> > > >>>
>> > > >>> I left some comments on EJ's doc.  I actually have a lot of
>> comments
>> > on
>> > > >>> the
>> > > >>> REST API design, I only posted some of them to start a discussion
>> > > >>> without overloading the doc.
>> > > >>>
>> > > >>> Overall, I believe EJ's proposal should also allow permission
>> > > assignments
>> > > >>> on tags that JB proposed (eventually). We just need to clearly
>> define
>> > > the
>> > > >>> persistence expectations for looking up related tags on the read
>> > path.
>> > > >>>
>> > > >>> We should probably specify whether and how tags are exposed to
>> > > >>> authorizers
>> > > >>> (OPA, Ranger). I imagine people will want to use them in external
>> > > policy
>> > > >>> engines the moment the feature is available.
>> > > >>>
>> > > >>> On the persistence side, I believe it would be nice to define
>> actual
>> > > java
>> > > >>> interfaces (perhaps in pseudo code) to allow easier review from
>> the
>> > > NoSQL
>> > > >>> persistence perspective (also commented in the doc).
>> > > >>>
>> > > >>> Cheers,
>> > > >>> Dmitri.
>> > > >>>
>> > > >>> On Thu, Jul 30, 2026 at 12:39 AM Jean-Baptiste Onofré <
>> > [email protected]
>> > > >
>> > > >>> wrote:
>> > > >>>
>> > > >>> > Hi EJ
>> > > >>> >
>> > > >>> > Thanks for starting this discussion.
>> > > >>> >
>> > > >>> > For the record, here's my initial proposal about tagging:
>> > > >>> >
>> https://lists.apache.org/thread/nmqmmjfmocfllb71fcmyp9syc9gyn820
>> > > >>> >
>> > > >>> > At that time, only Dmitri replied :)
>> > > >>> > So, I would be happy to work with you on this, as I still have
>> the
>> > > PoC
>> > > >>> > I created for my initial proposal.
>> > > >>> >
>> > > >>> > I will try to join the scheduled meeting (no guarantee).
>> > > >>> >
>> > > >>> > Regards
>> > > >>> > JB
>> > > >>> >
>> > > >>> > On Fri, Jul 17, 2026 at 6:53 AM EJ Wang <
>> > > >>> [email protected]>
>> > > >>> > wrote:
>> > > >>> > >
>> > > >>> > > Hi folks,
>> > > >>> > >
>> > > >>> > > I have prepared a Google Doc
>> > > >>> > > <
>> > > >>> >
>> > > >>>
>> > >
>> >
>> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
>> > > >>> > >
>> > > >>> > > for the Polaris tag spec proposal.
>> > > >>> > >
>> > > >>> > > The goal is simple: add a native tag model to Polaris so users
>> > can
>> > > >>> > classify
>> > > >>> > > catalog objects, read those classifications back, and find
>> > objects
>> > > by
>> > > >>> > tag.
>> > > >>> > >
>> > > >>> > > The proposal covers:
>> > > >>> > > * tag definitions as catalog-scoped Polaris entities
>> > > >>> > > * tag assignments on catalogs, namespaces, table-like objects,
>> > and
>> > > >>> > columns
>> > > >>> > > * allowed values on tag definitions
>> > > >>> > > * direct and inherited tag reads
>> > > >>> > > * direct by-tag lookup
>> > > >>> > > * the durable model behind the API
>> > > >>> > > * how this compares with the existing Polaris Policy API (tag
>> > > design
>> > > >>> > > referenced policy heavily, given their pattern similarity)
>> > > >>> > >
>> > > >>> > > Please take a look and leave comments in the doc. Let me know
>> > WDYT!
>> > > >>> > >
>> > > >>> > > I would also like to discuss this in the July 23 community
>> sync.
>> > A
>> > > >>> > separate
>> > > >>> > > dedicated review meeting will be scheduled separately, likely
>> > > within
>> > > >>> the
>> > > >>> > > next two weeks.
>> > > >>> > >
>> > > >>> > > Thanks,
>> > > >>> > > -ej
>> > > >>> >
>> > > >>>
>> > > >>
>> > >
>> >
>>
>

Reply via email to