Hi folks,

A quick update on the Tag work. All four planned PRs are now open:

   - PR1: API contract <https://github.com/apache/polaris/pull/5366>,
   updated following review.
   - PR2: Definition CRUD <https://github.com/apache/polaris/pull/5391>,
   Draft.
   - PR3: Assignment writes and storage
   <https://github.com/apache/polaris/pull/5469>, Draft.
   - PR4: Reads, inheritance, and reverse lookup
   <https://github.com/apache/polaris/pull/5595>, Draft.

PR4 completes the proposed implementation slices with direct and effective
tag reads, namespace inheritance, and lookup of objects assigned a tag. It
includes pagination and bounds on scan work. The PRs remain stacked and are
intended to merge in order.

Thanks Dmitri for confirming the API root and pagination direction here.
PR1 retains the current catalog root while broader API layout discussions
continue separately. Pagination is the default, with an explicit
pagination=false option subject to finite response-size and work limits.
That option returns the complete result or an error, never a silently
truncated result.

The design doc
<https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit>
remains the reference for the full proposal.

For PR1, please flag any remaining API contract concerns before we move
toward merging it. Early feedback on the implementation PRs is also
welcome, especially persistence consistency, authorization, and bounded
read behavior.

Thanks for the reviews so far!

-ej


On Wed, 23 Sep 2026 16:49:39 -0400, Dmitri Bourlatchkov <[email protected]>
wrote:

Hi EJ, On pagination: > [...] I'd still like to preserve an explicit
pagination=false option for > users accustomed to IRC's full-result mode
[...] > That option remains subject to finite server limits on response
size and > work. It returns the complete result or an error, never a
silently > truncated result. Requests exceeding those limits must use
pagination. This sounds reasonable to me. Cheers, Dmitri. On Mon, Sep 21,
2026 at 7:34 PM EJ Wang <[email protected]> wrote: > Hi
Dmitri, all, > > I've updated the design doc, and the corresponding PR
changes are in > progress: > >
https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit
> > On the API root, I see the benefits of independent routing and
versioning, > and agree that build wiring shouldn’t decide the paths. Would
you be > comfortable letting PR1 proceed under the current catalog root
while we > continue discussing the broader API layout? There are still
implementation > PRs ahead, so we can revisit this before release. If a
later decision > requires migration, we could introduce the new root while
retaining the > current paths for compatibility, accepting the maintenance
cost. > > On pagination, the doc now defaults to paginated responses, with
a > server-selected page size when omitted and client-provided sizes
treated as > hints. I understand that your oversized-response workaround
was > specifically for IRC compatibility. Tags has no existing-client >
compatibility requirement, but I'd still like to preserve an explicit >
pagination=false option for users accustomed to IRC's full-result mode as >
they adapt their workflows to paginated defaults. > > That option remains
subject to finite server limits on response size and > work. It returns the
complete result or an error, never a silently > truncated result. Requests
exceeding those limits must use pagination. > Would this bounded opt-out be
acceptable for Tags? > > Two corrections to my earlier replies to Robert
and Prithvi: reverse lookup > now requires per-target property-read checks,
and detach-all requires the > same atomic visible result from every
implementation supporting the API. > Physical cleanup may follow, but the
spec no longer defines a 501 fallback > for that guarantee. > > -ej > > On
Fri, Sep 18, 2026 at 1:31 PM Dmitri Bourlatchkov <[email protected]> >
wrote: > > > Hi EJ, > > > > > 1. *Base URL*: *existing catalog root* vs.
*independent Tag root* > > > > Forwarding my preference for a separate URI
base path (option B). I > believe > > I commented about that in GH too. > >
> > While each endpoint in the proposed Tags API is distinct from endpoints
> > currently under /api/catalog/polaris/v1, I do not think interleaving >
> endpoints from conceptually different APIs under the same URI base is a >
> good idea. > > > > If API under /api/catalog/polaris/v1 were formulated
in a way to allow > > exencibility, then adding tags there would be
natural. However, I do not > > think existing endpoints under that base URI
were designed for > > extencibility. The form very narrow purpose API.
Therefore, I believe it > is > > preferable to put tags under a separate
root. > > > > A smaller point is API evolution. Tags are defined in a
separate Open API > > YAML file. That scopes the API down and naturally
isolates it in terms of > > API evolution. > > > > Making changes to the
tags API will require careful consideration of the > > other APIs under the
same base URI to ensure no overlaps. > > > > That said, I do not feel too
strongly about the base URI. > > > > > 2. *Pagination*: *opt-in* vs.
*always-on* > > > > Forcing servers to provide unpaginated responses for
potentially large > > datasets is a DoS / overload risk, IMHO. More
in-depth discussion on this > > is in [1]. > > > > I would not want the
Polaris API specs to repeat that guideline from the > > IRC spec as I think
it is fundamentally flawed. > > > > I previously suggested [2] failing
large responses only as a means for > > maintaining IRC spec compatibility
in the IRC API. > > > > So I propose pagination to be "always on" in the
API spec. So, all > clients > > should be prepared to handle paginated
responses. > > > > However, pagination on the server side can be
implemented in phases, if > it > > helps with code-level PRs. > > > > [1]
https://lists.apache.org/thread/k81ptyktdbdf8gynncgk3o04mqt85zyk > > > >
[2] https://lists.apache.org/thread/ntn71oh7g0kkf1wdhskh8t986gdwh1p3 > > >
> Cheers, > > Dmitri > > > > On Tue, Sep 15, 2026 at 6:46 PM EJ Wang <
[email protected]> > > wrote: > > > > > Hi folks, > > > > > >
Following Dmitri's review of PR1, I'd like community input on two > choices
> > > before updating the spec. > > > > > > 1. *Base URL*: *existing
catalog root* vs. *independent Tag root* > > > > > > *A. Keep
/api/catalog/polaris/v1/{prefix}/.* Follow the existing policy > > and > >
> generic-table APIs, and discuss the broader extension API layout > > >
separately. > > > > > > *B. Move to /api/tags/v1/{prefix}/.* Give Tags
independent versioning > and > > > routing from its first release. Other
APIs would not need to move in > this > > > PR. > > > > > > *My preference
is A*, revising my earlier agreement to B in the PR. I > see > > > the
benefits of an independent root, but currently favor consistency > with > >
> the existing layout over introducing another pattern for Tags alone. > >
> > > > 2. *Pagination*: *opt-in* vs. *always-on* > > > > > > *A. Opt-in*.
Omitting pageToken requests all results in one response. > An > > > empty
token starts pagination. This follows Iceberg's convention and > lets > > >
simple clients avoid implementing pagination. > > > *B. Always-on*.
Omitting pagination parameters returns a default-sized > > > first page.
Clients must follow continuation tokens to retrieve all > > > results. This
is Dmitri's proposal and lets servers bound response > sizes. > > > > > >
*I'd like to explore A with an explicit safeguard*: oversized > unpaginated
> > > requests could fail with a documented error, never silently truncate
> > > results. That adds a limit to the full-result behavior, so it would >
need > > to > > > be part of the contract. > > > > > > The main concern is
reverse lookup, where a tag may reference many > > objects. > > > Internal
batching would not bound the total response size or request > > > duration.
The argument for A is client simplicity and convention > > > consistency,
not compatibility with existing Tag clients. > > > > > > Which option do
you favor for each? For pagination, would the > > > explicit-error
safeguard make A acceptable, or is B preferable from the > > > start? > > >
> > > Thanks, > > > -ej > > > > > > On Fri, Sep 11, 2026 at 7:55 AM Dmitri
Bourlatchkov <[email protected]> > > > wrote: > > > > > > > Hi All, > > > >
> > > > I posted a comment about tags API URI paths that may be of interest
> to > > a > > > > few people. Re-posting here for awareness: > > > > > > >
> https://github.com/apache/polaris/pull/5366#discussion_r3980198150 > > >
> > > > > Cheers, > > > > Dmitri. > > > > > > > > On Thu, Sep 10, 2026 at
1:21 AM EJ Wang < > > [email protected]> > > > > wrote: > > >
> > > > > > Hi folks, > > > > > > > > > > A quick update on the Tag work.
PR1-3 and the design doc have been > > > > updated > > > > > following the
discussion. > > > > > > > > > > *Current status*: > > > > > PR1: API
contract: ready for review. > > > > > PR2: Definition CRUD: open as Draft,
ready for review if PR1 LGTY. > > > > > *NEW! *PR3: Assignment writes and
storage: open as Draft, ready for > > > > review > > > > > if PR2 LGTY. > >
> > > PR4: Reads, inheritance, and reverse lookup: planned. > > > > > > > >
> > *More on PR3*: > > > > > PR3 adds assign/unassign for catalogs,
namespaces, tables, and > > > top-level > > > > > Iceberg columns, plus
atomic detach-all. It includes JDBC and > > in-memory > > > > > support.
Reads remain in PR4. > > > > > The PR is stacked on PR2. Its description
links the assignment-only > > > > commit > > > > > for focused review. > >
> > > > > > > > *Asks*: > > > > > For PR3, I’d especially appreciate
feedback on the persistence > > > > interfaces, > > > > > their impact on
external implementations, and the detach-all > > guarantee. > > > > > The
updated design doc also includes the reverse-lookup flow and a > > > > >
discussion of the deferred shared encoding work. > > > > > > > > > >
Thanks, > > > > > -ej > > > > > > > > > > On Wed, Sep 9, 2026 at 3:09 PM EJ
Wang < > > [email protected] > > > > > > > > > wrote: > > > >
> > > > > > > Thanks Robert, Prithvi, and Dmitri. I’ve updated the spec doc
> > > > > > < > > > > > > > > > > > > > > >
https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0#heading=h.jx650zq2bd2g
> > > > > > > > > > with > > > > > > the contract clarifications (see
below) and the follow-up > > discussion > > > > with > > > > > > Dmitri.
The corresponding PR updates are underway. > > > > > > > > > > > > *For
encoding*, my proposal is to retain Iceberg’s namespace > query > > > > > >
convention in v1, with explicit supported-name, encoding, and > > >
decoding > > > > > > rules in Part 1 section 5.2. This retains the known
separator and > > > > ingress > > > > > > limitations. Since the issue also
affects existing > Iceberg/Polaris > > > > APIs, > > > > > > I’d address a
replacement codec in a separate issue/PR. Part 3 > > > section > > > > >
7.12 > > > > > > records alternatives and spike results to start that
discussion. > > > > > > > > > > > > *Version tokens* are opaque strings in
both responses and update > > > > > > requests. Clients return the token
unchanged, and stale updates > > > return > > > > > 409. > > > > > > The
check must cover every supported definition-write path, while > > > > >
backends > > > > > > choose their revision mechanism. > > > > > > > > > > >
> *For detach-all*, the definition and assignments must disappear > > > >
together > > > > > > through Tag reads, or nothing changes. Physical
cleanup may > follow. > > > An > > > > > > implementation unable to provide
that guarantee returns 501 after > > > > > > authorization and before
changing visible state. > > > > > > > > > > > > *Column assignments* use
table identity and column identifier, > > using > > > > > > Iceberg field
id for Iceberg tables. Renames preserve assignments > > > when > > > > > >
identity is preserved, while same-name replacements do not > inherit > > >
> them. > > > > > V1 > > > > > > remains limited to top-level Iceberg
columns. > > > > > > > > > > > > Does this scope split work for you,
particularly keeping the > shared > > > > codec > > > > > > redesign
separate from Tag v1? I’d like to settle these contract > > > points > > >
> > > here before PR1 merges. > > > > > > > > > > > > -ej > > > > > > > > >
> > > On Wed, Sep 9, 2026 at 7:36 AM Dmitri Bourlatchkov < > >
[email protected] > > > > > > > > > > wrote: > > > > > > > > > > > >> Hi
All, > > > > > >> > > > > > >> (replying partially) > > > > > >> > > > > >
>> I very much support Robert's proposal for using a well-defined > and > >
> > > >> unambiguous format for namespaces. > > > > > >> > > > > > >> Many
Polaris APIs fall into following the IRC approach to > > namespace > > > >
> >> representation in query parameters. Yet, that approach has > >
multiple > > > > > issues > > > > > >> , which can be seen in Iceberg dev
ML / GH issues. > > > > > >> > > > > > >> I think Polaris should use a more
robust namespace > representation > > in > > > > its > > > > > >> native
APIs. > > > > > >> > > > > > >> Cheers, > > > > > >> Dmitri. > > > > > >> >
> > > > >> On Tue, Sep 8, 2026 at 6:41 AM Robert Stupp <[email protected]> > >
wrote: > > > > > >> > > > > > >> > Hi, > > > > > >> > > > > > > >> > I have
been thinking about what a full implementation would > > > require. > > > >
> >> > > > > > > >> > I like that the proposal separates definitions,
assignments, > and > > > > > >> effective > > > > > >> > reads. I have a
few API-contract questions that seem worth > > > resolving > > > > > >>
while > > > > > >> > the contract is still separate from the
implementation. > > > > > >> > > > > > > >> > > > > > > >> > First, I think
the target query parameters need a defined > > encoding > > > > for > > > >
> >> > identifier elements. > > > > > >> > > > > > > >> > This is not only
about the unit separator. Namespace elements > > and > > > > > object > > >
> > >> > names may themselves contain characters such as &, ?, =, +, or > >
%. > > > > > >> > Those must be preserved rather than interpreted as query
> syntax. > > > > > >> > > > > > > >> > Ordinary URI query-value encoding
handles those characters, > but > > it > > > > > does > > > > > >> not > >
> > > >> > solve the separate problem of representing the boundaries > >
between > > > > > >> multipart > > > > > >> > namespace elements. > > > > >
>> > > > > > > >> > I suggest defining a small, reversible
namespace-element > codec, > > > then > > > > > >> > applying ordinary URI
encoding to its complete output. > Nessie's > > > > > escaped > > > > > >>
> path > > > > > >> > representation is a useful precedent: it has an
unambiguous > > > element > > > > > >> > separator > > > > > >> > and
escape syntax, while avoiding control characters in the > > > > transport >
> > > > >> > representation. > > > > > >> > > > > > > >> > The contract
should specify the codec, its decoding failures, > > and > > > > > >> >
conformance examples. Client libraries should expose it rather > > > than >
> > > > >> > requiring > > > > > >> > every client to reproduce it. The
structured target used by > the > > > > write > > > > > >> APIs > > > > >
>> > would still be the clearest canonical representation; this > codec > >
> > would > > > > > >> make > > > > > >> > the GET form safe and
interoperable. > > > > > >> > > > > > > >> > > > > > > >> > Second, I think
the revision token should be opaque at the API > > > > > boundary. > > > >
> >> > > > > > > >> > The backend should be free to use a native row
revision, > commit > > > ID, > > > > > >> ETag, > > > > > >> > or > > > > >
>> > another conditional-write token. However, the contract should > > >
define > > > > > the > > > > > >> > observable precondition: the server
returns a token, and an > > update > > > > > >> succeeds > > > > > >> >
only if the client supplies the token for the current tag > > > >
definition. > > > > > >> > Otherwise the server returns a conflict. > > > >
> >> > > > > > > >> > That requires token matching semantics, but not an
integer > type, > > > an > > > > > >> initial > > > > > >> > value,
ordering, increment-by-one behavior, or history > > semantics. > > > A > >
> > > >> > catalog-wide commit token would also be valid, although it >
could > > > > > create > > > > > >> > avoidable conflicts for unrelated
changes. > > > > > >> > > > > > > >> > > > > > > >> > Third, the direct
reverse lookup is useful, but I would treat > it > > > as > > > > a > > > >
> >> > first-class, paginated relationship rather than a tag record > > > >
> containing > > > > > >> a > > > > > >> > collection of targets. > > > > >
>> > > > > > > >> > A common tag can legitimately be attached to a very
large > number > > > of > > > > > >> objects > > > > > >> > or columns. A
backend will normally need one forward access > path > > > for > > > > > >>
direct > > > > > >> > assignments by target, and one reverse access path by
> tag/value, > > > > with > > > > > >> > backend-specific partitioning or
sharding. Effective > assignments > > > > > should > > > > > >> > remain >
> > > > >> > computed from the target and its ancestors; materializing > >
> inherited > > > > > >> > assignments onto descendants would have very
different scaling > > > > > behavior. > > > > > >> > > > > > > >> > This
also affects detach-all. Deleting an unbounded number of > > > > >
assignment > > > > > >> > records atomically is not a portable primitive
for all > backends. > > > The > > > > > >> > contract > > > > > >> > should
distinguish observable deletion semantics from physical > > > > > cleanup,
> > > > > >> or > > > > > >> > state the backend capability required for a
synchronous > > detach-all > > > > > >> > operation. > > > > > >> > > > > >
> >> > > > > > > >> > Finally, I agree with the V1 boundary of top-level
Iceberg > > > columns, > > > > > but > > > > > >> I > > > > > >> > would >
> > > > >> > keep the core tag model independent of Iceberg. For Iceberg, >
the > > > > > durable > > > > > >> > column reference should be the field
ID, with a column name > used > > > > only > > > > > >> for > > > > > >> >
request-time resolution and display. Other table > implementations > > > >
could > > > > > >> opt > > > > > >> > in > > > > > >> > later once they
provide an equally stable field identity. This > > > > avoids > > > > > >>
> treating > > > > > >> > a name-based column mapping as a general
abstraction. > > > > > >> > > > > > > >> > > > > > > >> > None of this
requires tags to become authorization inputs in > V1. > > > It > > > > is >
> > > > >> > mainly > > > > > >> > about leaving the assignment and read
contract implementable > by > > > more > > > > > >> than > > > > > >> > one
> > > > > >> > persistence model when those slices arrive. > > > > > >> > >
> > > > >> > Thanks, > > > > > >> > Robert > > > > > >> > > > > > > >> > >
> > > > >> > On Fri, Aug 28, 2026 at 2:19 AM EJ Wang < > > > > >
[email protected] > > > > > >> > > > > > > >> > wrote: > > > >
> >> > > > > > > >> > > Hi folks, > > > > > >> > > > > > > > >> > > A quick
update on the Tag work. *Current status*: > > > > > >> > > - PR1: API
contract ( > > > https://github.com/apache/polaris/pull/5366 > > > > ): > >
> > > >> ready > > > > > >> > > for review > > > > > >> > > - *NEW! *PR2:
Definition CRUD ( > > > > > >> > https://github.com/apache/polaris/pull/5391
> > > > > >> > > ): > > > > > >> > > open as Draft, ready for review if PR1
LGTY > > > > > >> > > - PR3: Assignment writes and storage: planned > > > >
> >> > > - PR4: Reads, inheritance, and reverse lookup: planned > > > > >
>> > > > > > > > >> > > *More on PR2:* > > > > > >> > > - PR2 makes Tag
definitions usable through create, list, > load, > > > > > update, > > > >
> >> > > rename, and delete. It intentionally stops before > assignments, >
> > so > > > > > >> their > > > > > >> > > persistence model remains open
for the next slice. > > > > > >> > > - The PR is stacked on #5366 and will
be rebased once that > PR > > > > > merges. > > > > > >> > > > > > > > >> >
> *Asks:* > > > > > >> > > - For #5366, please call out any remaining API
contract > > > concerns. > > > > > For > > > > > >> > > #5391, I would
especially appreciate feedback on the slice > > > > boundary > > > > > >>
and > > > > > >> > the > > > > > >> > > decision to reuse the existing
entity persistence model. > > > > > >> > > - The design doc remains here: >
> > > > >> > > > > > > > >> > > > > > > > >> > > > > > > >> > > > > > > > >
> > > > > > >
https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
> > > > > >> > > > > > > > >> > > I’ll keep using this thread for new
delivery slices, > material > > > > status > > > > > >> > > changes, and
specific community asks. > > > > > >> > > > > > > > >> > > Thanks, > > > >
> >> > > -ej > > > > > >> > > > > > > > >> > > On Mon, Aug 24, 2026 at
5:04 PM EJ Wang < > > > > > >> [email protected]> > > > > > >>
> > wrote: > > > > > >> > > > > > > > >> > > > Hi folks, > > > > > >> > > >
> > > > > >> > > > Following up on this thread, I have opened a PR to land
> the > > > > public > > > > > >> API > > > > > >> > > > contract for Tags:
> > > https://github.com/apache/polaris/pull/5366 > > > > > >> > > > > > >
> > >> > > > The PR defines Tag management, assignment and > unassignment,
> > > > direct > > > > > >> and > > > > > >> > > > inherited reads, and
reverse lookup. V1 covers catalogs, > > > > > namespaces, > > > > > >> > >
> Iceberg and generic tables as whole objects, and top-level > > > >
Iceberg > > > > > >> > table > > > > > >> > > > columns. Views,
generic-table columns, nested fields, > > > > multi-value > > > > > >> > >
> assignments, and tag-based authorization are deferred. > > > > > >> > > >
> > > > > >> > > > I plan to deliver the capability through four PRs that >
merge > > > in > > > > > >> order: > > > > > >> > > the > > > > > >> > > >
API contract in this PR, Tag definition CRUD, assignment > > > writes > > >
> > and > > > > > >> > > > storage, then reads and reverse lookup. A
separate > follow-up > > > > will > > > > > >> add > > > > > >> > > >
grants on Tag resources to the management APIs. That grant > > > > surface
> > > > > >> is > > > > > >> > > > distinct from using Tags to control
access to tagged > > objects, > > > > > which > > > > > >> > > remains > >
> > > >> > > > outside v1. > > > > > >> > > > > > > > > >> > > > The
updated design doc is here: > > > > > >> > > > > > > > > >> > > > > > > >
>> > > > > > > >> > > > > > > > > > > > > > > >
https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
> > > > > >> > > > > > > > > >> > > > The PR is currently Draft while we
finish aligning on the > > > public > > > > > >> > > contract. > > > > > >>
> > > It is intended to merge as the first delivery slice, not > > > remain
> > > > > as a > > > > > >> > > > design-only artifact. Please call out any
remaining scope > or > > > > > >> contract > > > > > >> > > > concerns. If
the list is aligned, I will mark it ready for > > > > review. > > > > > >>
> > > > > > > > >> > > > Thanks, > > > > > >> > > > -ej > > > > > >> > > >
> > > > > >> > > > On Wed, Aug 12, 2026 at 2:05 PM EJ Wang < > > > > > >> >
[email protected]> > > > > > >> > > > wrote: > > > > > >> > >
> > > > > > >> > > >> Thanks Dmitri, these comments were very useful. > > >
> > >> > > >> > > > > > >> > > >> I went through the three areas you called
out and updated > > the > > > > > >> proposal > > > > > >> > > >>
accordingly. > > > > > >> > > >> > > > > > >> > > >> On the
permission/policy direction, *I agree the Tag > model > > > > should > > >
> > >> > leave > > > > > >> > > >> room for permissions or policies to
consume tags later*, > > > > > including > > > > > >> the > > > > > >> > >
>> direction JB proposed. I am keeping that outside the v1 > Tag > > > > >
>> contract, > > > > > >> > > >> though. In v1, tags classify resources;
they do not > > > themselves > > > > > >> grant > > > > > >> > or > > > > >
>> > > >> deny access. Polaris Policy looks like the closest > existing > >
> > > >> foundation > > > > > >> > > if > > > > > >> > > >> we later want a
portable tag-aware policy model, but I > > think > > > > that > > > > > >>
> > deserves > > > > > >> > > >> a separate proposal rather than baking
policy semantics > > into > > > > the > > > > > >> Tag > > > > > >> > > >>
storage model now. > > > > > >> > > >> > > > > > >> > > >> I also made the
authorizer path more explicit. *A future > > OPA, > > > > > >> Ranger, > >
> > > >> > or > > > > > >> > > >> other authorizer could receive the
target's complete > > > effective > > > > > >> tags as > > > > > >> > > >>
resource attributes*. The authorization path would > resolve > > > > those
> > > > > >> tags > > > > > >> > > >> internally, applying target-types,
inheritance, > > closest-wins, > > > > > >> > > grandfathered > > > > > >>
> > >> values, and the same coherent-read guarantees as the Tag > > API. >
> > > At > > > > > >> > > minimum, > > > > > >> > > >> the portable input
can include the tag definition ID, > > current > > > > > name, > > > > > >>
> and > > > > > >> > > >> selected value; provenance can be additional
context. If > > > > Polaris > > > > > >> > cannot > > > > > >> > > >>
resolve the complete effective state, authorization > should > > > fail > >
> > > >> closed > > > > > >> > > >> rather than treat the resource as
untagged. > > > > > >> > > >> > > > > > >> > > >> That also makes the
persistence expectation on the read > > path > > > > > >> clearer: > > > >
> >> > an > > > > > >> > > >> implementation needs to resolve the target
and relevant > > > > > ancestors, > > > > > >> > > obtain > > > > > >> > >
>> the applicable tag definitions and assignments, and > produce > > > one
> > > > > >> > coherent > > > > > >> > > >> effective result. *Those
observable semantics are the > > backend > > > > > >> > contract; > > > > >
>> > > >> the physical lookup/indexing strategy is not.* > > > > > >> > >
>> > > > > > >> > > >> On the Java interface suggestion, I added
Java-shaped > > records > > > > for > > > > > >> the > > > > > >> > > >>
durable logical model so the definition, target identity, > > and > > > > >
>> > assignment > > > > > >> > > >> shapes are easier to review from JDBC
and NoSQL > > > perspectives. I > > > > > >> > stopped > > > > > >> > > >>
short of proposing operation interfaces in pseudo-code, > > > though. > > >
> > My > > > > > >> > > current > > > > > >> > > >> thinking is that we
should first agree on the durable > facts > > > and > > > > > >> > required
> > > > > >> > > >> behavior, then design the actual persistence SPI around
> the > > > > needs > > > > > >> of > > > > > >> > the > > > > > >> > > >>
implementations. I did not want an illustrative interface > > in > > > >
this > > > > > >> > > design to > > > > > >> > > >> accidentally become the
persistence contract. > > > > > >> > > >> > > > > > >> > > >> So Part 2 now
separates the two intentionally: > > > > > >> > > >> > > > > > >> > > >>
*logical data + behavior/conformance requirements are > > > > specified; >
> > > > >> > > >> transaction, CAS, atomic batch, provider-native >
operations, > > > and > > > > > the > > > > > >> > > >> eventual Java SPI
remain implementation/design choices.* > > > > > >> > > >> > > > > > >> > >
>> Thanks again for the review, and definitely keep the > > comments > > >
> > >> coming > > > > > >> > :) > > > > > >> > > >> > > > > > >> > > >>
I've updated the doc, please check it out the latest and > > the > > > > >
>> greatest: > > > > > >> > > >> > > > > > >> > > >> > > > > > >> > > > > >
> > >> > > > > > > >> > > > > > > > > > > > > > > >
https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0
> > > > > >> > > >> > > > > > >> > > >> -ej > > > > > >> > > >> > > > > >
>> > > >> On Fri, Aug 7, 2026 at 3:39 PM Dmitri Bourlatchkov < > > > > > >>
[email protected]> > > > > > >> > > >> wrote: > > > > > >> > > >> > > > > >
>> > > >>> Hi EJ, JB, > > > > > >> > > >>> > > > > > >> > > >>> I left some
comments on EJ's doc. I actually have a lot > > of > > > > > >> comments >
> > > > >> > on > > > > > >> > > >>> the > > > > > >> > > >>> REST API
design, I only posted some of them to start a > > > > > discussion > > > >
> >> > > >>> without overloading the doc. > > > > > >> > > >>> > > > > > >>
> > >>> Overall, I believe EJ's proposal should also allow > > > permission
> > > > > >> > > assignments > > > > > >> > > >>> on tags that JB proposed
(eventually). We just need to > > > clearly > > > > > >> define > > > > >
>> > > the > > > > > >> > > >>> persistence expectations for looking up
related tags on > > the > > > > read > > > > > >> > path. > > > > > >> > >
>>> > > > > > >> > > >>> We should probably specify whether and how tags
are > > exposed > > > to > > > > > >> > > >>> authorizers > > > > > >> > >
>>> (OPA, Ranger). I imagine people will want to use them in > > > > >
external > > > > > >> > > policy > > > > > >> > > >>> engines the moment
the feature is available. > > > > > >> > > >>> > > > > > >> > > >>> On the
persistence side, I believe it would be nice to > > > define > > > > > >>
actual > > > > > >> > > java > > > > > >> > > >>> interfaces (perhaps in
pseudo code) to allow easier > review > > > > from > > > > > >> the > > > >
> >> > > NoSQL > > > > > >> > > >>> persistence perspective (also commented
in the doc). > > > > > >> > > >>> > > > > > >> > > >>> Cheers, > > > > > >>
> > >>> Dmitri. > > > > > >> > > >>> > > > > > >> > > >>> On Thu, Jul 30,
2026 at 12:39 AM Jean-Baptiste Onofré < > > > > > >> > [email protected] > >
> > > >> > > > > > > > > >> > > >>> wrote: > > > > > >> > > >>> > > > > >
>> > > >>> > Hi EJ > > > > > >> > > >>> > > > > > > >> > > >>> > Thanks for
starting this discussion. > > > > > >> > > >>> > > > > > > >> > > >>> > For
the record, here's my initial proposal about > > tagging: > > > > > >> > >
>>> > > > > > > >> >
https://lists.apache.org/thread/nmqmmjfmocfllb71fcmyp9syc9gyn820 > > > > >
>> > > >>> > > > > > > >> > > >>> > At that time, only Dmitri replied :) >
> > > > >> > > >>> > So, I would be happy to work with you on this, as I >
> still > > > > have > > > > > >> the > > > > > >> > > PoC > > > > > >> > >
>>> > I created for my initial proposal. > > > > > >> > > >>> > > > > > >
>> > > >>> > I will try to join the scheduled meeting (no > guarantee). > >
> > > >> > > >>> > > > > > > >> > > >>> > Regards > > > > > >> > > >>> > JB
> > > > > >> > > >>> > > > > > > >> > > >>> > On Fri, Jul 17, 2026 at
6:53 AM EJ Wang < > > > > > >> > > >>> [email protected]> > >
> > > >> > > >>> > wrote: > > > > > >> > > >>> > > > > > > > >> > > >>> > >
Hi folks, > > > > > >> > > >>> > > > > > > > >> > > >>> > > I have prepared
a Google Doc > > > > > >> > > >>> > > < > > > > > >> > > >>> > > > > > > >>
> > >>> > > > > > >> > > > > > > > >> > > > > > > >> > > > > > > > > > > >
> > > >
https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
> > > > > >> > > >>> > > > > > > > >> > > >>> > > for the Polaris tag spec
proposal. > > > > > >> > > >>> > > > > > > > >> > > >>> > > The goal is
simple: add a native tag model to > Polaris > > so > > > > > users > > > >
> >> > can > > > > > >> > > >>> > classify > > > > > >> > > >>> > > catalog
objects, read those classifications back, > and > > > find > > > > > >> >
objects > > > > > >> > > by > > > > > >> > > >>> > tag. > > > > > >> > >
>>> > > > > > > > >> > > >>> > > The proposal covers: > > > > > >> > > >>>
> > * tag definitions as catalog-scoped Polaris entities > > > > > >> > >
>>> > > * tag assignments on catalogs, namespaces, > table-like > > > > >
objects, > > > > > >> > and > > > > > >> > > >>> > columns > > > > > >> > >
>>> > > * allowed values on tag definitions > > > > > >> > > >>> > > *
direct and inherited tag reads > > > > > >> > > >>> > > * direct by-tag
lookup > > > > > >> > > >>> > > * the durable model behind the API > > > >
> >> > > >>> > > * how this compares with the existing Polaris Policy > >
API > > > > > (tag > > > > > >> > > design > > > > > >> > > >>> > >
referenced policy heavily, given their pattern > > > similarity) > > > > >
>> > > >>> > > > > > > > >> > > >>> > > Please take a look and leave
comments in the doc. > Let > > me > > > > > know > > > > > >> > WDYT! > > >
> > >> > > >>> > > > > > > > >> > > >>> > > I would also like to discuss
this in the July 23 > > > community > > > > > >> sync. > > > > > >> > A > >
> > > >> > > >>> > separate > > > > > >> > > >>> > > dedicated review
meeting will be scheduled > separately, > > > > > likely > > > > > >> > >
within > > > > > >> > > >>> the > > > > > >> > > >>> > > next two weeks. >
> > > > >> > > >>> > > > > > > > >> > > >>> > > Thanks, > > > > > >> > >
>>> > > -ej > > > > > >> > > >>> > > > > > > >> > > >>> > > > > > >> > > >>
> > > > > >> > > > > > > > >> > > > > > > >> > > > > > > > > > > > > > > >
> > > > > >

Reply via email to