Hi,

On Wed, Jul 22, 2026 at 12:12 AM huaxin gao <[email protected]> wrote:
>
> Thanks Xiening, that's a fair concern, and I agree the fast update/delete 
> need doesn't go away.
>
> I'd frame it less as shifting the burden and more as changing where it's 
> paid. Equality deletes push the cost onto every read: each reader re-joins 
> the delete files against candidate rows, on every query, for the life of the 
> table. Resolving position once (at write time, or once during background 
> conversion) pays that cost a single time and makes all later reads cheap, an 
> O(1) DV check. And the streaming writer stays cheap either way: background 
> conversion (or write-time DV resolution via the index) moves the 
> position-resolution cost off the hot write path, so the write stays cheap 
> while every downstream reader gets fast, position-based deletes instead of 
> re-joining equality-delete files on each query.
>
> On index cost: I agree keeping a key -> position index current isn't free, 
> but the evidence so far is that it's manageable. Max's ConvertEqualityDeletes 
> already maintains a persistent RocksDB PK index incrementally at streaming 
> scale, not rebuilt each cycle, so this is running in practice, not just in 
> theory.

I'm a little confused about the index part:
1. If we already have a key -> position index, why can't the reader
use it to resolve equality deletes into row positions at read time,
assuming the equality_ids match the index key?
2. Regarding the RocksDB-based primary-key index mentioned above, how
would we maintain such an index efficiently in cloud object storage?
Is the idea that the index remains local to a specific engine, rather
than being stored and shared as part of the Iceberg table?

>
> On tooling and adoption: tooling is central to the proposal, not an 
> afterthought. Background conversion already lets writers keep working 
> unchanged, they write equality deletes as usual, and the conversion job 
> resolves them to DVs using its own internal index. The persistent key-lookup 
> index is the shared tooling that makes write-time elimination practical and 
> engine-agnostic. Because forbidding is a V4-table property, streaming-upsert 
> workloads that don't yet have a write-time DV path can keep running on V3 
> (writing equality deletes, with background conversion keeping reads fast), 
> and move to V4 once their engine can emit DVs directly. Non-streaming 
> workloads can adopt V4 right away. So it shouldn't force anyone into a worse 
> position or block adoption.
>
> Finally, removing equality deletes isn't only about read cost. They also 
> block CDC, row lineage, and incremental maintenance of indexes and 
> materialized views, so it's less "shift the burden" and more "unblock 
> features that equality deletes currently make impossible."
>
> Thanks,
> Huaxin
>
> On Mon, Jul 20, 2026 at 3:20 PM Xiening Dai <[email protected]> wrote:
>>
>> Hi Huaxin,
>>
>> Thanks for bringing this up. Equality delete is indeed a pain point we have 
>> seen in many customer use cases.
>>
>> That been said the scenario of fast update/delete still exists no matter 
>> which technology or table spec we choose. The proposal is just going to 
>> shift the burden from the reader to the writer. To achieve fast 
>> update/delete, customer can build index structure like you mentioned, but 
>> building such index and keeping it up to date all time can be very expensive 
>> too (especially given that we are tackling the fast update/delete scenario). 
>> So although it feels like a right direction as we are saying that we don't 
>> want to handle this complexity on the table spec, the underlying problem is 
>> not solved. Without a good alternative or tooling support, i am afraid this 
>> could become an adoption issue for v4 going forward.
>>
>> On 2026/07/17 17:13:55 huaxin gao wrote:
>> > Thanks all for the discussion. I've thought this over, and I'd like to
>> > change my view to forbidding equality-delete writes for V4 tables, rather
>> > than the softer "deprecated but permitted."
>> >
>> > My earlier hesitation was about gating V4 on a write-time DV implementation
>> > that isn't built yet. I think that concern goes away once we separate the
>> > spec decision from engine adoption:
>> >
>> > - Forbidding equality deletes is a V4 format decision. It defines what a V4
>> > table allows; it does not require every engine to have the write-time DV
>> > path on day one.
>> > - Workloads that still rely on streaming upserts can stay on V3 until their
>> > engine's write-time path is ready, then adopt V4. Readers continue to
>> > support equality deletes for existing V2 and V3 tables.
>> > - So we don't need the write-time implementation finished to forbid
>> > equality deletes in V4. What we need is a clear, credible path, and I think
>> > we have it: Flink's ConvertEqualityDeletes already maintains a PK index in
>> > Flink state and does key -> position -> DV today, and the next step is to
>> > persist that index into Iceberg so any engine can resolve positions and
>> > emit DVs directly at write time.
>> >
>> > This also addresses Max's concern that "deprecated but permitted" is too
>> > soft. Forbidding for V4 gives engines a real incentive to move, while
>> > keeping existing tables fully readable.
>> >
>> > On Xin's point, agreed that some perf data comparing DVs on V4 against
>> > equality deletes on V3 would be useful to have as we go.
>> >
>> > Thanks,
>> > Huaxin
>> >
>> > On Thu, Jul 16, 2026 at 8:43 AM Xin Huang via dev <[email protected]>
>> > wrote:
>> >
>> > > Conceptually +1 as equality deletes really complicates the format and
>> > > implementation.
>> > >
>> > > However given the concern is around performance side. Is there a way to
>> > > make the decision making more data driven — having some benchmarking on
>> > > perf comparison between dv+single file commit in v4 vs equality delete in
>> > > v3 could help making a call.
>> > >
>> > > Thanks
>> > > Xin
>> > >
>> > > On Wed, Jul 15, 2026 at 11:01 PM Maximilian Michels <[email protected]>
>> > > wrote:
>> > >
>> > >> I understood "deprecate equality deletes" as not forbidding engines to
>> > >> write them, but rather discouraging them. IMHO this is long overdue,
>> > >> but it is also a very soft transition. Perhaps too soft, because it
>> > >> doesn't give engines who write them a real incentive to stop writing
>> > >> equality deletes.
>> > >>
>> > >> Engines will likely be quicker to move away from writing equality
>> > >> deletes if we disallow writing them in V4. Regardless, we will have to
>> > >> support reading equality deletes for V2 and V3 tables.
>> > >>
>> > >> I'm leaning more towards removing equality deletes for V4 tables, but
>> > >> I would like to hear what others think.
>> > >>
>> > >> -Max
>> > >>
>> > >>
>> > >> On Wed, Jul 15, 2026 at 10:03 PM Steven Wu <[email protected]> wrote:
>> > >> >
>> > >> > > So I'd separate two things: deprecating in V4 (signal and direction,
>> > >> safe to do now) versus forbidding equality-delete writes (gated on the
>> > >> engine-agnostic path being ready). I'm only proposing the first for V4.
>> > >> >
>> > >> > I thought we wanted to forbid equality-delete writes for v4 tables,
>> > >> which would really simplify the v4 adaptive metadata tree along with 
>> > >> other
>> > >> benefits that Huaxin already outlined,
>> > >> >
>> > >> > > deprecation path in v4
>> > >> >
>> > >> > I heard the Kafka connector in the Iceberg repo doesn't produce
>> > >> equality deletes. We would need to migrate the Flink sink to leverage 
>> > >> the
>> > >> index to produce DVs only in v4.
>> > >> >
>> > >> > On Wed, Jul 15, 2026 at 12:13 PM huaxin gao <[email protected]>
>> > >> wrote:
>> > >> >>
>> > >> >> Thanks Max and Manu.
>> > >> >>
>> > >> >> Max, thanks for the added detail. It's a good point that the index in
>> > >> ConvertEqualityDeletes is persisted in Flink state (RocksDB) and updated
>> > >> incrementally. That strengthens the case, since it shows the key to
>> > >> position index is already durable, just scoped to Flink today. And your
>> > >> closing point is exactly the plan I have in mind: deprecate equality
>> > >> deletes in V4, and once the spec has an index, persist that PK index 
>> > >> into
>> > >> Iceberg so it can be shared across engines.
>> > >> >>
>> > >> >> Manu, good question on how this works in practice. To be clear, I'm
>> > >> proposing "deprecated but permitted," not removal. In V4, writers
>> > >> (including Kafka Connect) could keep emitting equality deletes and 
>> > >> readers
>> > >> would keep applying them. Deprecation just declares deletion vectors the
>> > >> going-forward mechanism and stops new investment in equality deletes. 
>> > >> So V4
>> > >> would not be gated on the index landing.
>> > >> >>
>> > >> >> On the cleanup path today: ConvertEqualityDeletes runs against the
>> > >> table, not a specific writer, so a Kafka Connect pipeline can already be
>> > >> cleaned up. It keeps writing equality deletes, and a standalone Flink
>> > >> ConvertEqualityDeletes job converts them to DVs. The one friction is 
>> > >> that
>> > >> the conversion runtime is Flink today, so a Kafka-only shop would have 
>> > >> to
>> > >> run Flink just for maintenance. A Spark action would be a natural 
>> > >> follow-up
>> > >> here, since Spark is the usual Iceberg maintenance engine and most batch
>> > >> shops already run it.
>> > >> >>
>> > >> >> Longer term, the next step is to persist the key to position index
>> > >> into Iceberg. Once it's shared, engines can look up positions and write 
>> > >> DVs
>> > >> directly at write time, so a writer can stop producing equality deletes
>> > >> entirely, and neither the Flink job nor a Spark action needs to rebuild 
>> > >> the
>> > >> index each run. Each engine (Kafka Connect, Spark, Flink) would adopt
>> > >> write-time DVs on its own schedule; until then it keeps writing equality
>> > >> deletes and relies on background conversion. So V4 deprecation doesn't
>> > >> require re-implementing every writer up front.
>> > >> >>
>> > >> >> So I'd separate two things: deprecating in V4 (signal and direction,
>> > >> safe to do now) versus forbidding equality-delete writes (gated on the
>> > >> engine-agnostic path being ready). I'm only proposing the first for V4.
>> > >> >>
>> > >> >> Thanks,
>> > >> >> Huaxin
>> > >> >>
>> > >> >> On Wed, Jul 15, 2026 at 3:10 AM Manu Zhang <[email protected]>
>> > >> wrote:
>> > >> >>>
>> > >> >>> Hi Huaxin,
>> > >> >>>
>> > >> >>> +1 for deprecating equality deletes, but how would this deprecation
>> > >> work practically in V4?
>> > >> >>> As Max pointed out, we still lack an engine-agnostic solution for
>> > >> streaming use cases. For example, how would we handle equality deletes
>> > >> written by Kafka Connect?
>> > >> >>> While the index proposal looks promising, I don't see a clear path
>> > >> for deprecating equality deletes in V4 before that index work actually
>> > >> lands.
>> > >> >>>
>> > >> >>> Thanks,
>> > >> >>> Manu
>> > >> >>>
>> > >> >>>
>> > >> >>> On Wed, Jul 15, 2026 at 5:30 PM Maximilian Michels <[email protected]>
>> > >> wrote:
>> > >> >>>>
>> > >> >>>> Hi Huaxin,
>> > >> >>>>
>> > >> >>>> Thanks for reviving the discussion on deprecating equality deletes.
>> > >> >>>> Equality deletes are the number one pain for streaming use cases.
>> > >> Many
>> > >> >>>> users give up when they see the merge-on-read costs, or they build
>> > >> >>>> custom solutions which move them further away from core Iceberg. 
>> > >> >>>> That
>> > >> >>>> said, we've made great progress since the initial conversation in
>> > >> >>>> 2024.
>> > >> >>>>
>> > >> >>>> Just to add what you said: The index we maintain in
>> > >> >>>> ConvertEqualityDeletes is not ephemeral. The index is persisted in
>> > >> >>>> Flink's managed state (RocksDB). It is continuously updated as new
>> > >> >>>> data arrives and checkpointed periodically. However, even though 
>> > >> >>>> the
>> > >> >>>> conversion works for data written by any engine, we currently 
>> > >> >>>> require
>> > >> >>>> Flink for the conversion itself. Storing the index directly in
>> > >> Iceberg
>> > >> >>>> and enabling all engines access would be the next logical step
>> > >> towards
>> > >> >>>> a fully engine-agnostic solution.
>> > >> >>>>
>> > >> >>>> The reality is that we don't yet have a working solution to avoid
>> > >> >>>> writing equality deletes across all engines, but given the recent
>> > >> >>>> progress, the proposed plan seems realistic. So +1 for deprecating
>> > >> >>>> equality deletes in V4.
>> > >> >>>>
>> > >> >>>> Cheers,
>> > >> >>>> Max
>> > >> >>>>
>> > >> >>>>
>> > >> >>>>
>> > >> >>>> On Tue, Jul 14, 2026 at 3:25 AM huaxin gao <[email protected]>
>> > >> wrote:
>> > >> >>>> >
>> > >> >>>> > Hi all,
>> > >> >>>> >
>> > >> >>>> > I'd like to restart the conversation about deprecating equality
>> > >> deletes, now in the context of the V4 spec.
>> > >> >>>> >
>> > >> >>>> > Background
>> > >> >>>> >
>> > >> >>>> > This isn't a new idea. Russell proposed deprecating equality
>> > >> deletes in V3 and removing them from the spec in V4, back in October 
>> > >> 2024
>> > >> in "[DISCUSS] - Deprecate Equality Deletes". The main blocker at the 
>> > >> time
>> > >> was that equality deletes served real use cases (especially Flink 
>> > >> streaming
>> > >> upserts) with no efficient alternative. Two developments since then make
>> > >> the V4 removal worth acting on now.
>> > >> >>>> >
>> > >> >>>> > Why equality deletes are costly
>> > >> >>>> >
>> > >> >>>> > Equality deletes are cheap to write but expensive to read: a
>> > >> reader must load the equality-delete files and join them against every
>> > >> candidate row in the delete's sequence-number range. Positional deletes
>> > >> skip that per-row join by marking exact positions, so they have always 
>> > >> read
>> > >> faster, and V3 deletion vectors make them faster still, one compact 
>> > >> bitmap
>> > >> per data file, applied by an O(1) position check, instead of V2's many
>> > >> position-delete files. So equality deletes' only real edge is the cheap
>> > >> write, and both background conversion and a write-time key-lookup index 
>> > >> can
>> > >> recover that.
>> > >> >>>> >
>> > >> >>>> > Beyond performance
>> > >> >>>> >
>> > >> >>>> > Equality deletes also block other features. CDC and row lineage
>> > >> are effectively impossible while they are in use, because the true 
>> > >> state of
>> > >> the table can only be determined with a full scan. That same property 
>> > >> means
>> > >> differential structures such as materialized views and secondary indexes
>> > >> have to be fully rebuilt whenever an equality delete is added, rather 
>> > >> than
>> > >> maintained incrementally. So removing equality deletes is close to a
>> > >> prerequisite for the index work to stay incrementally maintainable.
>> > >> >>>> >
>> > >> >>>> > Evidence the alternatives are practical
>> > >> >>>> >
>> > >> >>>> > 1. Converting equality deletes to DVs works today. Max Michels'
>> > >> ConvertEqualityDeletes maintenance task (16831, 16844, 16858, 16874, 
>> > >> 16889,
>> > >> 16948) rewrites equality deletes into deletion vectors as a background
>> > >> Flink job: the writer keeps appending equality deletes to a staging 
>> > >> branch,
>> > >> and the task converts them to DVs on the target branch so reads apply
>> > >> deletes by position. Notably, the task resolves each delete to a 
>> > >> position
>> > >> using a primary-key index that it builds and maintains inside the job,
>> > >> demonstrating the full "key -> position -> DV" path end to end.
>> > >> >>>> >
>> > >> >>>> > 2. A persistent key-lookup index removes the need to write them 
>> > >> >>>> > at
>> > >> all. The secondary index spec we're working on (#16961) includes a
>> > >> key-lookup index mapping a key to its data file and row position. This 
>> > >> is
>> > >> essentially the persistent, catalog-managed form of the index Max's task
>> > >> builds ephemerally. With it, a writer can resolve positions at write 
>> > >> time
>> > >> and emit DVs directly, without ever producing an equality delete.
>> > >> >>>> >
>> > >> >>>> > How these two efforts fit together
>> > >> >>>> >
>> > >> >>>> > They're complementary, and they cover the two things we need to
>> > >> deprecate equality deletes:
>> > >> >>>> >
>> > >> >>>> > Migration (existing data): ConvertEqualityDeletes cleans up 
>> > >> >>>> > tables
>> > >> that already contain equality deletes, and supports writers that still 
>> > >> emit
>> > >> them, converting them to DVs in the background.
>> > >> >>>> > Going forward (new writes): the persistent key-lookup index lets
>> > >> writers skip equality deletes entirely by looking up positions directly.
>> > >> >>>> > The connection is that Max's task already proves the core
>> > >> mechanism (resolve key -> position, write a DV); it just rebuilds a
>> > >> throwaway index each cycle. A durable, shared index both enables 
>> > >> write-time
>> > >> elimination and removes that rebuild cost from the conversion path.
>> > >> >>>> >
>> > >> >>>> >
>> > >> >>>> > Proposal
>> > >> >>>> >
>> > >> >>>> > I propose that we deprecate equality deletes in V4. The blocker
>> > >> from 2024 was the lack of a viable alternative, and we now have the 
>> > >> pieces:
>> > >> background conversion to DVs works today, and the key-lookup index 
>> > >> gives us
>> > >> a path to eliminating them at write time. Deletion vectors should be the
>> > >> going-forward mechanism for row-level deletes and upserts, produced by
>> > >> background conversion now and directly by writers once the index is
>> > >> available. Readers would continue to support equality deletes for 
>> > >> backward
>> > >> compatibility with existing V2/V3 tables.
>> > >> >>>> >
>> > >> >>>> > Migration path
>> > >> >>>> >
>> > >> >>>> > Existing tables keep working; readers continue to apply equality
>> > >> deletes.
>> > >> >>>> > ConvertEqualityDeletes (Flink) rewrites existing equality deletes
>> > >> into DVs so tables can be cleared of them over time.
>> > >> >>>> >
>> > >> >>>> >
>> > >> >>>> > I'd love people's thoughts, especially from those running large
>> > >> streaming-upsert workloads.
>> > >> >>>> >
>> > >> >>>> > Thanks,
>> > >> >>>> > Huaxin
>> > >>
>> > >
>> >



-- 
Regards
Junwang Zhao

Reply via email to