Hi folks,
Following Dmitri's review of PR1, I'd like community input on two choices
before updating the spec.
1. *Base URL*: *existing catalog root* vs. *independent Tag root*
*A. Keep /api/catalog/polaris/v1/{prefix}/.* Follow the existing policy and
generic-table APIs, and discuss the broader extension API layout separately.
*B. Move to /api/tags/v1/{prefix}/.* Give Tags independent versioning and
routing from its first release. Other APIs would not need to move in this
PR.
*My preference is A*, revising my earlier agreement to B in the PR. I see
the benefits of an independent root, but currently favor consistency with
the existing layout over introducing another pattern for Tags alone.
2. *Pagination*: *opt-in* vs. *always-on*
*A. Opt-in*. Omitting pageToken requests all results in one response. An
empty token starts pagination. This follows Iceberg's convention and lets
simple clients avoid implementing pagination.
*B. Always-on*. Omitting pagination parameters returns a default-sized
first page. Clients must follow continuation tokens to retrieve all
results. This is Dmitri's proposal and lets servers bound response sizes.
*I'd like to explore A with an explicit safeguard*: oversized unpaginated
requests could fail with a documented error, never silently truncate
results. That adds a limit to the full-result behavior, so it would need to
be part of the contract.
The main concern is reverse lookup, where a tag may reference many objects.
Internal batching would not bound the total response size or request
duration. The argument for A is client simplicity and convention
consistency, not compatibility with existing Tag clients.
Which option do you favor for each? For pagination, would the
explicit-error safeguard make A acceptable, or is B preferable from the
start?
Thanks,
-ej
On Fri, Sep 11, 2026 at 7:55 AM Dmitri Bourlatchkov <[email protected]>
wrote:
> Hi All,
>
> I posted a comment about tags API URI paths that may be of interest to a
> few people. Re-posting here for awareness:
>
> https://github.com/apache/polaris/pull/5366#discussion_r3980198150
>
> Cheers,
> Dmitri.
>
> On Thu, Sep 10, 2026 at 1:21 AM EJ Wang <[email protected]>
> wrote:
>
> > Hi folks,
> >
> > A quick update on the Tag work. PR1-3 and the design doc have been
> updated
> > following the discussion.
> >
> > *Current status*:
> > PR1: API contract: ready for review.
> > PR2: Definition CRUD: open as Draft, ready for review if PR1 LGTY.
> > *NEW! *PR3: Assignment writes and storage: open as Draft, ready for
> review
> > if PR2 LGTY.
> > PR4: Reads, inheritance, and reverse lookup: planned.
> >
> > *More on PR3*:
> > PR3 adds assign/unassign for catalogs, namespaces, tables, and top-level
> > Iceberg columns, plus atomic detach-all. It includes JDBC and in-memory
> > support. Reads remain in PR4.
> > The PR is stacked on PR2. Its description links the assignment-only
> commit
> > for focused review.
> >
> > *Asks*:
> > For PR3, I’d especially appreciate feedback on the persistence
> interfaces,
> > their impact on external implementations, and the detach-all guarantee.
> > The updated design doc also includes the reverse-lookup flow and a
> > discussion of the deferred shared encoding work.
> >
> > Thanks,
> > -ej
> >
> > On Wed, Sep 9, 2026 at 3:09 PM EJ Wang <[email protected]>
> > wrote:
> >
> > > Thanks Robert, Prithvi, and Dmitri. I’ve updated the spec doc
> > > <
> >
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0#heading=h.jx650zq2bd2g
> >
> > with
> > > the contract clarifications (see below) and the follow-up discussion
> with
> > > Dmitri. The corresponding PR updates are underway.
> > >
> > > *For encoding*, my proposal is to retain Iceberg’s namespace query
> > > convention in v1, with explicit supported-name, encoding, and decoding
> > > rules in Part 1 section 5.2. This retains the known separator and
> ingress
> > > limitations. Since the issue also affects existing Iceberg/Polaris
> APIs,
> > > I’d address a replacement codec in a separate issue/PR. Part 3 section
> > 7.12
> > > records alternatives and spike results to start that discussion.
> > >
> > > *Version tokens* are opaque strings in both responses and update
> > > requests. Clients return the token unchanged, and stale updates return
> > 409.
> > > The check must cover every supported definition-write path, while
> > backends
> > > choose their revision mechanism.
> > >
> > > *For detach-all*, the definition and assignments must disappear
> together
> > > through Tag reads, or nothing changes. Physical cleanup may follow. An
> > > implementation unable to provide that guarantee returns 501 after
> > > authorization and before changing visible state.
> > >
> > > *Column assignments* use table identity and column identifier, using
> > > Iceberg field id for Iceberg tables. Renames preserve assignments when
> > > identity is preserved, while same-name replacements do not inherit
> them.
> > V1
> > > remains limited to top-level Iceberg columns.
> > >
> > > Does this scope split work for you, particularly keeping the shared
> codec
> > > redesign separate from Tag v1? I’d like to settle these contract points
> > > here before PR1 merges.
> > >
> > > -ej
> > >
> > > On Wed, Sep 9, 2026 at 7:36 AM Dmitri Bourlatchkov <[email protected]>
> > > wrote:
> > >
> > >> Hi All,
> > >>
> > >> (replying partially)
> > >>
> > >> I very much support Robert's proposal for using a well-defined and
> > >> unambiguous format for namespaces.
> > >>
> > >> Many Polaris APIs fall into following the IRC approach to namespace
> > >> representation in query parameters. Yet, that approach has multiple
> > issues
> > >> , which can be seen in Iceberg dev ML / GH issues.
> > >>
> > >> I think Polaris should use a more robust namespace representation in
> its
> > >> native APIs.
> > >>
> > >> Cheers,
> > >> Dmitri.
> > >>
> > >> On Tue, Sep 8, 2026 at 6:41 AM Robert Stupp <[email protected]> wrote:
> > >>
> > >> > Hi,
> > >> >
> > >> > I have been thinking about what a full implementation would require.
> > >> >
> > >> > I like that the proposal separates definitions, assignments, and
> > >> effective
> > >> > reads. I have a few API-contract questions that seem worth resolving
> > >> while
> > >> > the contract is still separate from the implementation.
> > >> >
> > >> >
> > >> > First, I think the target query parameters need a defined encoding
> for
> > >> > identifier elements.
> > >> >
> > >> > This is not only about the unit separator. Namespace elements and
> > object
> > >> > names may themselves contain characters such as &, ?, =, +, or %.
> > >> > Those must be preserved rather than interpreted as query syntax.
> > >> >
> > >> > Ordinary URI query-value encoding handles those characters, but it
> > does
> > >> not
> > >> > solve the separate problem of representing the boundaries between
> > >> multipart
> > >> > namespace elements.
> > >> >
> > >> > I suggest defining a small, reversible namespace-element codec, then
> > >> > applying ordinary URI encoding to its complete output. Nessie's
> > escaped
> > >> > path
> > >> > representation is a useful precedent: it has an unambiguous element
> > >> > separator
> > >> > and escape syntax, while avoiding control characters in the
> transport
> > >> > representation.
> > >> >
> > >> > The contract should specify the codec, its decoding failures, and
> > >> > conformance examples. Client libraries should expose it rather than
> > >> > requiring
> > >> > every client to reproduce it. The structured target used by the
> write
> > >> APIs
> > >> > would still be the clearest canonical representation; this codec
> would
> > >> make
> > >> > the GET form safe and interoperable.
> > >> >
> > >> >
> > >> > Second, I think the revision token should be opaque at the API
> > boundary.
> > >> >
> > >> > The backend should be free to use a native row revision, commit ID,
> > >> ETag,
> > >> > or
> > >> > another conditional-write token. However, the contract should define
> > the
> > >> > observable precondition: the server returns a token, and an update
> > >> succeeds
> > >> > only if the client supplies the token for the current tag
> definition.
> > >> > Otherwise the server returns a conflict.
> > >> >
> > >> > That requires token matching semantics, but not an integer type, an
> > >> initial
> > >> > value, ordering, increment-by-one behavior, or history semantics. A
> > >> > catalog-wide commit token would also be valid, although it could
> > create
> > >> > avoidable conflicts for unrelated changes.
> > >> >
> > >> >
> > >> > Third, the direct reverse lookup is useful, but I would treat it as
> a
> > >> > first-class, paginated relationship rather than a tag record
> > containing
> > >> a
> > >> > collection of targets.
> > >> >
> > >> > A common tag can legitimately be attached to a very large number of
> > >> objects
> > >> > or columns. A backend will normally need one forward access path for
> > >> direct
> > >> > assignments by target, and one reverse access path by tag/value,
> with
> > >> > backend-specific partitioning or sharding. Effective assignments
> > should
> > >> > remain
> > >> > computed from the target and its ancestors; materializing inherited
> > >> > assignments onto descendants would have very different scaling
> > behavior.
> > >> >
> > >> > This also affects detach-all. Deleting an unbounded number of
> > assignment
> > >> > records atomically is not a portable primitive for all backends. The
> > >> > contract
> > >> > should distinguish observable deletion semantics from physical
> > cleanup,
> > >> or
> > >> > state the backend capability required for a synchronous detach-all
> > >> > operation.
> > >> >
> > >> >
> > >> > Finally, I agree with the V1 boundary of top-level Iceberg columns,
> > but
> > >> I
> > >> > would
> > >> > keep the core tag model independent of Iceberg. For Iceberg, the
> > durable
> > >> > column reference should be the field ID, with a column name used
> only
> > >> for
> > >> > request-time resolution and display. Other table implementations
> could
> > >> opt
> > >> > in
> > >> > later once they provide an equally stable field identity. This
> avoids
> > >> > treating
> > >> > a name-based column mapping as a general abstraction.
> > >> >
> > >> >
> > >> > None of this requires tags to become authorization inputs in V1. It
> is
> > >> > mainly
> > >> > about leaving the assignment and read contract implementable by more
> > >> than
> > >> > one
> > >> > persistence model when those slices arrive.
> > >> >
> > >> > Thanks,
> > >> > Robert
> > >> >
> > >> >
> > >> > On Fri, Aug 28, 2026 at 2:19 AM EJ Wang <
> > [email protected]
> > >> >
> > >> > wrote:
> > >> >
> > >> > > Hi folks,
> > >> > >
> > >> > > A quick update on the Tag work. *Current status*:
> > >> > > - PR1: API contract (https://github.com/apache/polaris/pull/5366
> ):
> > >> ready
> > >> > > for review
> > >> > > - *NEW! *PR2: Definition CRUD (
> > >> > https://github.com/apache/polaris/pull/5391
> > >> > > ):
> > >> > > open as Draft, ready for review if PR1 LGTY
> > >> > > - PR3: Assignment writes and storage: planned
> > >> > > - PR4: Reads, inheritance, and reverse lookup: planned
> > >> > >
> > >> > > *More on PR2:*
> > >> > > - PR2 makes Tag definitions usable through create, list, load,
> > update,
> > >> > > rename, and delete. It intentionally stops before assignments, so
> > >> their
> > >> > > persistence model remains open for the next slice.
> > >> > > - The PR is stacked on #5366 and will be rebased once that PR
> > merges.
> > >> > >
> > >> > > *Asks:*
> > >> > > - For #5366, please call out any remaining API contract concerns.
> > For
> > >> > > #5391, I would especially appreciate feedback on the slice
> boundary
> > >> and
> > >> > the
> > >> > > decision to reuse the existing entity persistence model.
> > >> > > - The design doc remains here:
> > >> > >
> > >> > >
> > >> >
> > >>
> >
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
> > >> > >
> > >> > > I’ll keep using this thread for new delivery slices, material
> status
> > >> > > changes, and specific community asks.
> > >> > >
> > >> > > Thanks,
> > >> > > -ej
> > >> > >
> > >> > > On Mon, Aug 24, 2026 at 5:04 PM EJ Wang <
> > >> [email protected]>
> > >> > > wrote:
> > >> > >
> > >> > > > Hi folks,
> > >> > > >
> > >> > > > Following up on this thread, I have opened a PR to land the
> public
> > >> API
> > >> > > > contract for Tags: https://github.com/apache/polaris/pull/5366
> > >> > > >
> > >> > > > The PR defines Tag management, assignment and unassignment,
> direct
> > >> and
> > >> > > > inherited reads, and reverse lookup. V1 covers catalogs,
> > namespaces,
> > >> > > > Iceberg and generic tables as whole objects, and top-level
> Iceberg
> > >> > table
> > >> > > > columns. Views, generic-table columns, nested fields,
> multi-value
> > >> > > > assignments, and tag-based authorization are deferred.
> > >> > > >
> > >> > > > I plan to deliver the capability through four PRs that merge in
> > >> order:
> > >> > > the
> > >> > > > API contract in this PR, Tag definition CRUD, assignment writes
> > and
> > >> > > > storage, then reads and reverse lookup. A separate follow-up
> will
> > >> add
> > >> > > > grants on Tag resources to the management APIs. That grant
> surface
> > >> is
> > >> > > > distinct from using Tags to control access to tagged objects,
> > which
> > >> > > remains
> > >> > > > outside v1.
> > >> > > >
> > >> > > > The updated design doc is here:
> > >> > > >
> > >> > >
> > >> >
> > >>
> >
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
> > >> > > >
> > >> > > > The PR is currently Draft while we finish aligning on the public
> > >> > > contract.
> > >> > > > It is intended to merge as the first delivery slice, not remain
> > as a
> > >> > > > design-only artifact. Please call out any remaining scope or
> > >> contract
> > >> > > > concerns. If the list is aligned, I will mark it ready for
> review.
> > >> > > >
> > >> > > > Thanks,
> > >> > > > -ej
> > >> > > >
> > >> > > > On Wed, Aug 12, 2026 at 2:05 PM EJ Wang <
> > >> > [email protected]>
> > >> > > > wrote:
> > >> > > >
> > >> > > >> Thanks Dmitri, these comments were very useful.
> > >> > > >>
> > >> > > >> I went through the three areas you called out and updated the
> > >> proposal
> > >> > > >> accordingly.
> > >> > > >>
> > >> > > >> On the permission/policy direction, *I agree the Tag model
> should
> > >> > leave
> > >> > > >> room for permissions or policies to consume tags later*,
> > including
> > >> the
> > >> > > >> direction JB proposed. I am keeping that outside the v1 Tag
> > >> contract,
> > >> > > >> though. In v1, tags classify resources; they do not themselves
> > >> grant
> > >> > or
> > >> > > >> deny access. Polaris Policy looks like the closest existing
> > >> foundation
> > >> > > if
> > >> > > >> we later want a portable tag-aware policy model, but I think
> that
> > >> > > deserves
> > >> > > >> a separate proposal rather than baking policy semantics into
> the
> > >> Tag
> > >> > > >> storage model now.
> > >> > > >>
> > >> > > >> I also made the authorizer path more explicit. *A future OPA,
> > >> Ranger,
> > >> > or
> > >> > > >> other authorizer could receive the target's complete effective
> > >> tags as
> > >> > > >> resource attributes*. The authorization path would resolve
> those
> > >> tags
> > >> > > >> internally, applying target-types, inheritance, closest-wins,
> > >> > > grandfathered
> > >> > > >> values, and the same coherent-read guarantees as the Tag API.
> At
> > >> > > minimum,
> > >> > > >> the portable input can include the tag definition ID, current
> > name,
> > >> > and
> > >> > > >> selected value; provenance can be additional context. If
> Polaris
> > >> > cannot
> > >> > > >> resolve the complete effective state, authorization should fail
> > >> closed
> > >> > > >> rather than treat the resource as untagged.
> > >> > > >>
> > >> > > >> That also makes the persistence expectation on the read path
> > >> clearer:
> > >> > an
> > >> > > >> implementation needs to resolve the target and relevant
> > ancestors,
> > >> > > obtain
> > >> > > >> the applicable tag definitions and assignments, and produce one
> > >> > coherent
> > >> > > >> effective result. *Those observable semantics are the backend
> > >> > contract;
> > >> > > >> the physical lookup/indexing strategy is not.*
> > >> > > >>
> > >> > > >> On the Java interface suggestion, I added Java-shaped records
> for
> > >> the
> > >> > > >> durable logical model so the definition, target identity, and
> > >> > assignment
> > >> > > >> shapes are easier to review from JDBC and NoSQL perspectives. I
> > >> > stopped
> > >> > > >> short of proposing operation interfaces in pseudo-code, though.
> > My
> > >> > > current
> > >> > > >> thinking is that we should first agree on the durable facts and
> > >> > required
> > >> > > >> behavior, then design the actual persistence SPI around the
> needs
> > >> of
> > >> > the
> > >> > > >> implementations. I did not want an illustrative interface in
> this
> > >> > > design to
> > >> > > >> accidentally become the persistence contract.
> > >> > > >>
> > >> > > >> So Part 2 now separates the two intentionally:
> > >> > > >>
> > >> > > >> *logical data + behavior/conformance requirements are
> specified;
> > >> > > >> transaction, CAS, atomic batch, provider-native operations, and
> > the
> > >> > > >> eventual Java SPI remain implementation/design choices.*
> > >> > > >>
> > >> > > >> Thanks again for the review, and definitely keep the comments
> > >> coming
> > >> > :)
> > >> > > >>
> > >> > > >> I've updated the doc, please check it out the latest and the
> > >> greatest:
> > >> > > >>
> > >> > > >>
> > >> > >
> > >> >
> > >>
> >
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0
> > >> > > >>
> > >> > > >> -ej
> > >> > > >>
> > >> > > >> On Fri, Aug 7, 2026 at 3:39 PM Dmitri Bourlatchkov <
> > >> [email protected]>
> > >> > > >> wrote:
> > >> > > >>
> > >> > > >>> Hi EJ, JB,
> > >> > > >>>
> > >> > > >>> I left some comments on EJ's doc. I actually have a lot of
> > >> comments
> > >> > on
> > >> > > >>> the
> > >> > > >>> REST API design, I only posted some of them to start a
> > discussion
> > >> > > >>> without overloading the doc.
> > >> > > >>>
> > >> > > >>> Overall, I believe EJ's proposal should also allow permission
> > >> > > assignments
> > >> > > >>> on tags that JB proposed (eventually). We just need to clearly
> > >> define
> > >> > > the
> > >> > > >>> persistence expectations for looking up related tags on the
> read
> > >> > path.
> > >> > > >>>
> > >> > > >>> We should probably specify whether and how tags are exposed to
> > >> > > >>> authorizers
> > >> > > >>> (OPA, Ranger). I imagine people will want to use them in
> > external
> > >> > > policy
> > >> > > >>> engines the moment the feature is available.
> > >> > > >>>
> > >> > > >>> On the persistence side, I believe it would be nice to define
> > >> actual
> > >> > > java
> > >> > > >>> interfaces (perhaps in pseudo code) to allow easier review
> from
> > >> the
> > >> > > NoSQL
> > >> > > >>> persistence perspective (also commented in the doc).
> > >> > > >>>
> > >> > > >>> Cheers,
> > >> > > >>> Dmitri.
> > >> > > >>>
> > >> > > >>> On Thu, Jul 30, 2026 at 12:39 AM Jean-Baptiste Onofré <
> > >> > [email protected]
> > >> > > >
> > >> > > >>> wrote:
> > >> > > >>>
> > >> > > >>> > Hi EJ
> > >> > > >>> >
> > >> > > >>> > Thanks for starting this discussion.
> > >> > > >>> >
> > >> > > >>> > For the record, here's my initial proposal about tagging:
> > >> > > >>> >
> > >> https://lists.apache.org/thread/nmqmmjfmocfllb71fcmyp9syc9gyn820
> > >> > > >>> >
> > >> > > >>> > At that time, only Dmitri replied :)
> > >> > > >>> > So, I would be happy to work with you on this, as I still
> have
> > >> the
> > >> > > PoC
> > >> > > >>> > I created for my initial proposal.
> > >> > > >>> >
> > >> > > >>> > I will try to join the scheduled meeting (no guarantee).
> > >> > > >>> >
> > >> > > >>> > Regards
> > >> > > >>> > JB
> > >> > > >>> >
> > >> > > >>> > On Fri, Jul 17, 2026 at 6:53 AM EJ Wang <
> > >> > > >>> [email protected]>
> > >> > > >>> > wrote:
> > >> > > >>> > >
> > >> > > >>> > > Hi folks,
> > >> > > >>> > >
> > >> > > >>> > > I have prepared a Google Doc
> > >> > > >>> > > <
> > >> > > >>> >
> > >> > > >>>
> > >> > >
> > >> >
> > >>
> >
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
> > >> > > >>> > >
> > >> > > >>> > > for the Polaris tag spec proposal.
> > >> > > >>> > >
> > >> > > >>> > > The goal is simple: add a native tag model to Polaris so
> > users
> > >> > can
> > >> > > >>> > classify
> > >> > > >>> > > catalog objects, read those classifications back, and find
> > >> > objects
> > >> > > by
> > >> > > >>> > tag.
> > >> > > >>> > >
> > >> > > >>> > > The proposal covers:
> > >> > > >>> > > * tag definitions as catalog-scoped Polaris entities
> > >> > > >>> > > * tag assignments on catalogs, namespaces, table-like
> > objects,
> > >> > and
> > >> > > >>> > columns
> > >> > > >>> > > * allowed values on tag definitions
> > >> > > >>> > > * direct and inherited tag reads
> > >> > > >>> > > * direct by-tag lookup
> > >> > > >>> > > * the durable model behind the API
> > >> > > >>> > > * how this compares with the existing Polaris Policy API
> > (tag
> > >> > > design
> > >> > > >>> > > referenced policy heavily, given their pattern similarity)
> > >> > > >>> > >
> > >> > > >>> > > Please take a look and leave comments in the doc. Let me
> > know
> > >> > WDYT!
> > >> > > >>> > >
> > >> > > >>> > > I would also like to discuss this in the July 23 community
> > >> sync.
> > >> > A
> > >> > > >>> > separate
> > >> > > >>> > > dedicated review meeting will be scheduled separately,
> > likely
> > >> > > within
> > >> > > >>> the
> > >> > > >>> > > next two weeks.
> > >> > > >>> > >
> > >> > > >>> > > Thanks,
> > >> > > >>> > > -ej
> > >> > > >>> >
> > >> > > >>>
> > >> > > >>
> > >> > >
> > >> >
> > >>
> > >
> >
>