Hi All,

Re: Default location reuse with soft delete

I do believe we cannot reuse the name-based location because it will open
data confidentiality attack vectors.

Therefore, we must purge the data before allowing the soft-deleted name to
be reused by a different entity.

I'll try to make another PR review round soon, but I cannot promise any ETL
:)

Cheers,
Dmitri.


On Sat, Sep 19, 2026 at 4:43 AM Prithvi S <[email protected]>
wrote:

> Hi JB, Dmitri,
>
> Thanks for the replies, and thanks Dmitri for looking at the draft PR.
>
> JB, on eviction: Phase 1 is lazy on namespace access, not a scheduled
> sweep.
> When TABLE_SOFT_DELETE_ENABLED is false (the default), expiration is a
> no-op: a catalog config check only. Catalogs that keep hard DROP do not
> scan held tables.
> When it is enabled, the next list, create, or drop-namespace in that
> namespace lists Iceberg tables under it (live and dropped) and permanently
> deletes those whose toPurgeTimestamp has passed. Catalog state is always
> removed then. File cleanup is a separate flag and uses the existing
> ENTITY_CLEANUP_SCHEDULER. There is no dedicated expiration task, and no
> purge-now API, in Phase 1 :)
> So the cost you asked about only applies to catalogs that opt in. A
> namespace with many held tables then pays a full Iceberg-table listing on
> those operations. I documented that on TABLE_SOFT_DELETE_HOLD_PERIOD in the
> draft PR. A scheduled sweep can wait until we know whether lazy-on-access
> is enough.
>
> Dmitri also noted on the PR that leaving files in place after expire, then
> freeing the name, can let a recreated table see leftover data until unique
> locations are enforced. The DISCUSS default was file cleanup off so expire
> matches DROP without purge. If Phase 1 should always enqueue entity-cleanup
> when a soft-deleted table is permanently deleted, I can change the draft.
>
> Thanks,
> Prithvi S
>
> On Sat, Sep 19, 2026 at 5:54 AM Dmitri Bourlatchkov <[email protected]>
> wrote:
>
> > Hi Prithvi,
> >
> > The proposed two phase approach sounds reasonable to me. I'll try and
> > review the PR ASAP.
> >
> > Cheers,
> > Dmitri.
> >
> > On Thu, Sep 17, 2026 at 10:25 PM Prithvi S <[email protected]>
> > wrote:
> >
> > > Hi all,
> > >
> > > I'd like to start a discussion on opt-in table soft-delete, hold
> period,
> > > and later undrop: https://github.com/apache/polaris/issues/5054
> > >
> > > Polaris DROP is hard from the catalog's point of view: the active name
> > > disappears, and there is no hold period or undrop. Operators who need a
> > > recoverable drop (compliance / legal hold, Nessie-to-Polaris migration)
> > > invent workarounds that Iceberg REST clients can bypass.
> > >
> > > The change is opt-in and default-off, so catalogs that want today's
> hard
> > > DROP keep it.
> > >
> > > Thanks to Dmitri for looking at this earlier. His suggestion was that
> > name
> > > reuse after soft-delete should be blocked until tables have unique
> > > locations (PR #4606), because shared locations plus a recreated table
> > > interact badly with staged-create credential refresh.
> > >
> > > #4606 has merged, but unique locations are still off by default,
> existing
> > > tables are unchanged, and client-specified locations are still allowed.
> > So
> > > "all tables have unique locations" is not true today, and I do not want
> > > #5054 to wait on that migration.
> > >
> > > The original GitHub issue freed the table name as soon as the table was
> > > soft-deleted, so a drop-and-recreate ETL job could create the same
> > > identifier again during the hold. I now think that is the wrong default
> > for
> > > an opt-in hold/recover feature:
> > >   * A legal-hold table should not share its identifier with a new live
> > > table.
> > >   * Undrop is unambiguous if the name is still reserved.
> > >   * Drop-and-recreate ETL is unchanged on catalogs that leave the
> feature
> > > off (the default). Catalogs that enable it can DROP PURGE when they
> truly
> > > want the name back immediately.
> > >
> > > Phase 1 (default off): Catalog-scoped, INTERNAL Iceberg tables only.
> > > Disabled catalogs keep
> > > today's DROP behavior.
> > >   * DROP without purge -> soft-delete: entity retained, metadata
> location
> > > retained, data/metadata files not deleted.
> > >   * Soft-deleted tables are not listable or loadable on normal Iceberg
> > REST
> > > paths.
> > >   * The identifier stays reserved for the hold period (default 7 days).
> > > CREATE / register with the same name fails until permanent delete or a
> > > later undrop.
> > >   * DROP PURGE on a live table remains immediate permanent delete,
> using
> > > the existing cleanup path.
> > >   * After the hold, the next list/create/drop-namespace in that
> namespace
> > > permanently deletes catalog state. File cleanup is opt-in (default off)
> > and
> > > uses today's ENTITY_CLEANUP_SCHEDULER.
> > >   * No new Iceberg REST undrop. No change to unique-location defaults.
> > >   * No new privileges in Phase 1: soft-drop uses TABLE_DROP; purge
> keeps
> > > the current drop-with-purge authorization.
> > >
> > > Persistence: in-place dropped entity with dropTimestamp /
> > toPurgeTimestamp,
> > > filtered out of the active resolve/list path. Do not revive the old
> > > entitiesDropped SPI.
> > >
> > > Phase 2: List soft-deleted tables and undrop, with new privileges. Name
> > > reuse after soft-delete only if/when a catalog actually has unique
> > > locations for every table.
> > >
> > > out of scope for this thread:
> > >   * Staged-create credential refresh
> > >   * Flipping DEFAULT_UNIQUE_TABLE_LOCATION_ENABLED default
> > >
> > > questions:
> > >   1. Is reserving the identifier during the hold acceptable for Phase
> 1,
> > > instead of freeing the table name on soft-delete?
> > >   2. Is it correct that DROP PURGE stays "destroy now", and only DROP
> > > without purge enters the hold?
> > >
> > > I opened a draft PR for Phase 1:
> > > https://github.com/apache/polaris/pull/5544
> > >
> > > Thanks,
> > > Prithvi S
> > >
> >
>

Reply via email to