I also agree that we don't need to fully replace equality deletes before
moving forward. Demonstrating that they are not required is a reasonable
standard, since v3 continues to work. We don't want to tightly couple Flink
equality delete work with v4.

On Fri, Jul 24, 2026 at 10:15 AM Steven Wu <[email protected]> wrote:

> I explicitly support disallowing the writing of new equality deletes in
> v4. Reading existing equality deletes would still be supported in the
> reference implementation as others mentioned.
>
> I echo Dan's comment on the benchmark result Viquar shared. The
> *IcebergSourceParquetEqDeleteBenchmark* (from the Java repo) doesn't
> fully capture the read performance problem of equality deletes, as it only
> applies one equality file to the data file. Note that equality delete files
> apply to all previously added data files (in the same partition scope or
> globally). In production workloads, each data file can have hundreds,
> thousands, or even more equality delete files attached to it, which makes
> both scan planning and file reading very expensive.
>
> I agree with Russell that as long as we see a clear alternative path for
> the streaming writer like Flink, we don't have to complete the replacement
> implementation before v4 certification. It would certainly be great to have
> though.
>
> On Fri, Jul 24, 2026 at 9:46 AM Ryan Blue <[email protected]> wrote:
>
>> > Do you mean we should allow a V3 table that contains equality delete to
>> upgrade to V4? Sound like you'd suggest the V4 table spec would still allow
>> the presence of equality delete. In  that case, there's no way to enforce
>> writer not to generate EQ deletes for V4.
>>
>> That's correct: I think that an upgrade can leave existing equality
>> deletes in place.
>>
>> I don't agree that there is no way to enforce writers not generating
>> equality deletes. If we don't allow equality deletes to be written, then we
>> don't need to support storing them in v4 manifest files. We would keep the
>> v3 delete manifests around until the delete files age off or are compacted
>> into the data files. Not being able to store new equality deletes in v4
>> manifests would be effective.
>>
>> Going further, I think what you're suggesting is that as long as we allow
>> v3 delete manifests, a client could produce delete files and write them to
>> a v3 manifest file. It's true that misbehaving clients could see this as a
>> hack and do it. But spec-compliant implementations would not allow this and
>> can easily check for it. The same goes for catalogs, so someone doing this
>> would need to control their client and their catalog. At that point, it's a
>> lot of work to maintain essentially a fork of the format. I don't think
>> anyone would do this, but if we find it's a problem we can always disallow
>> upgrading the next version (v5) if there are v3 delete manifests present.
>>
>> People can do things that aren't allowed by the format, but that
>> shouldn't stop us from moving forward when it is a healthy choice to drop
>> support for something that has been a persistent problem.
>>
>> Ryan
>>
>> On Thu, Jul 23, 2026 at 3:38 PM Xiening Dai <[email protected]> wrote:
>>
>>> > I think we can continue to support equality deletes on upgrade.
>>>
>>> Do you mean we should allow a V3 table that contains equality delete to
>>> upgrade to V4? Sound like you'd suggest the V4 table spec would still allow
>>> the presence of equality delete. In  that case, there's no way to enforce
>>> writer not to generate EQ deletes for V4.
>>>
>>> Or should we require the engine to perform a synchronized EQ delete
>>> re-write during the upgrade?
>>>
>>> On 2026/07/23 21:34:13 Ryan Blue wrote:
>>> > Also, I forgot to include how to handle existing tables. I think we can
>>> > continue to support equality deletes on upgrade. The main benefit of
>>> > getting rid of them would be in scan planning, where we have to run
>>> 2-phase
>>> > planning to match delete files to data files. But we can't get rid of
>>> > 2-phase planning yet because position deletes (DVs) also require
>>> 2-phase
>>> > planning.
>>> >
>>> > For v4, we should check for delete manifests and perform 2-phase
>>> planning
>>> > if there are any. Once all DVs are co-located with data files and
>>> equality
>>> > deletes are rewritten, there will be no delete manifests and we can
>>> skip it.
>>> >
>>> > On Thu, Jul 23, 2026 at 2:22 PM Ryan Blue <[email protected]> wrote:
>>> >
>>> > > I strongly support this proposal. We should deprecate equality
>>> deletes and
>>> > > not allow writing them in v4 tables.
>>> > >
>>> > > The Flink work shows that it is reasonable to maintain an index so
>>> the
>>> > > application can create DVs rather than using equality deletes. This
>>> is
>>> > > independent of the on-going index work where the index can be used
>>> for
>>> > > other purposes. I don't think moving to DVs is blocked by the index
>>> work,
>>> > > although it will be great when Flink can produce and share its index.
>>> > >
>>> > > Now that the index-based solution is demonstrated, I think that
>>> > > disallowing equality deletes in v4 is the right path forward.
>>> Equality
>>> > > deletes are much, much more expensive to apply than DVs. In the v3
>>> release,
>>> > > we moved from multiple position delete files to a single DV and I
>>> see this
>>> > > decision as very similar: we thought that table maintenance would be
>>> viable
>>> > > to mitigate the downside of having multiple delete files, but that
>>> just
>>> > > wasn't the case so we should move the responsibility. Instead of
>>> relying on
>>> > > async maintenance that didn't materialize, we need to make this a
>>> writer
>>> > > responsibility.
>>> > >
>>> > > I agree with the points that Huaxin wrote. Existing workloads can
>>> continue
>>> > > to use v3 until they are moved over to produce DVs. The DV-based
>>> write can
>>> > > happen in v3 tables so there is a clean upgrade path for the writer
>>> before
>>> > > upgrading a table.
>>> > >
>>> > > The thread also brought up Kafka Connect, where producing DVs is a
>>> bit
>>> > > harder because KC doesn't shuffle data. I don't think that this is a
>>> > > blocker for a few reasons:
>>> > > 1. The Apache Iceberg KC bundle deliberately does not support upsert
>>> > > because of the problems caused by equality deletes. We chose to limit
>>> > > functionality rather than make tables unusable at read time.
>>> > > 2. Anyone self-supporting their own KC build that writes equality
>>> deletes
>>> > > can continue to use v3 without issues
>>> > > 3. There are solutions for writing DVs from KC. I think the primary
>>> > > challenge is first repartitioning the data. Just because KC doesn't
>>> do this
>>> > > today doesn't mean we shouldn't move forward when moving this
>>> > > responsibility to writers is the right trade-off.
>>> > >
>>> > > Thanks for all your work on this, Huaxin and Max! I think this is
>>> going to
>>> > > be a major step forward.
>>> > >
>>> > > Ryan
>>> > >
>>> > > On Thu, Jul 23, 2026 at 2:28 AM Junwang Zhao <[email protected]>
>>> wrote:
>>> > >
>>> > >> Hi,
>>> > >>
>>> > >> On Wed, Jul 22, 2026 at 12:12 AM huaxin gao <[email protected]
>>> >
>>> > >> wrote:
>>> > >> >
>>> > >> > Thanks Xiening, that's a fair concern, and I agree the fast
>>> > >> update/delete need doesn't go away.
>>> > >> >
>>> > >> > I'd frame it less as shifting the burden and more as changing
>>> where
>>> > >> it's paid. Equality deletes push the cost onto every read: each
>>> reader
>>> > >> re-joins the delete files against candidate rows, on every query,
>>> for the
>>> > >> life of the table. Resolving position once (at write time, or once
>>> during
>>> > >> background conversion) pays that cost a single time and makes all
>>> later
>>> > >> reads cheap, an O(1) DV check. And the streaming writer stays cheap
>>> either
>>> > >> way: background conversion (or write-time DV resolution via the
>>> index)
>>> > >> moves the position-resolution cost off the hot write path, so the
>>> write
>>> > >> stays cheap while every downstream reader gets fast, position-based
>>> deletes
>>> > >> instead of re-joining equality-delete files on each query.
>>> > >> >
>>> > >> > On index cost: I agree keeping a key -> position index current
>>> isn't
>>> > >> free, but the evidence so far is that it's manageable. Max's
>>> > >> ConvertEqualityDeletes already maintains a persistent RocksDB PK
>>> index
>>> > >> incrementally at streaming scale, not rebuilt each cycle, so this is
>>> > >> running in practice, not just in theory.
>>> > >>
>>> > >> I'm a little confused about the index part:
>>> > >> 1. If we already have a key -> position index, why can't the reader
>>> > >> use it to resolve equality deletes into row positions at read time,
>>> > >> assuming the equality_ids match the index key?
>>> > >> 2. Regarding the RocksDB-based primary-key index mentioned above,
>>> how
>>> > >> would we maintain such an index efficiently in cloud object storage?
>>> > >> Is the idea that the index remains local to a specific engine,
>>> rather
>>> > >> than being stored and shared as part of the Iceberg table?
>>> > >>
>>> > >> >
>>> > >> > On tooling and adoption: tooling is central to the proposal, not
>>> an
>>> > >> afterthought. Background conversion already lets writers keep
>>> working
>>> > >> unchanged, they write equality deletes as usual, and the conversion
>>> job
>>> > >> resolves them to DVs using its own internal index. The persistent
>>> > >> key-lookup index is the shared tooling that makes write-time
>>> elimination
>>> > >> practical and engine-agnostic. Because forbidding is a V4-table
>>> property,
>>> > >> streaming-upsert workloads that don't yet have a write-time DV path
>>> can
>>> > >> keep running on V3 (writing equality deletes, with background
>>> conversion
>>> > >> keeping reads fast), and move to V4 once their engine can emit DVs
>>> > >> directly. Non-streaming workloads can adopt V4 right away. So it
>>> shouldn't
>>> > >> force anyone into a worse position or block adoption.
>>> > >> >
>>> > >> > Finally, removing equality deletes isn't only about read cost.
>>> They
>>> > >> also block CDC, row lineage, and incremental maintenance of indexes
>>> and
>>> > >> materialized views, so it's less "shift the burden" and more
>>> "unblock
>>> > >> features that equality deletes currently make impossible."
>>> > >> >
>>> > >> > Thanks,
>>> > >> > Huaxin
>>> > >> >
>>> > >> > On Mon, Jul 20, 2026 at 3:20 PM Xiening Dai <[email protected]>
>>> wrote:
>>> > >> >>
>>> > >> >> Hi Huaxin,
>>> > >> >>
>>> > >> >> Thanks for bringing this up. Equality delete is indeed a pain
>>> point we
>>> > >> have seen in many customer use cases.
>>> > >> >>
>>> > >> >> That been said the scenario of fast update/delete still exists no
>>> > >> matter which technology or table spec we choose. The proposal is
>>> just going
>>> > >> to shift the burden from the reader to the writer. To achieve fast
>>> > >> update/delete, customer can build index structure like you
>>> mentioned, but
>>> > >> building such index and keeping it up to date all time can be very
>>> > >> expensive too (especially given that we are tackling the fast
>>> update/delete
>>> > >> scenario). So although it feels like a right direction as we are
>>> saying
>>> > >> that we don't want to handle this complexity on the table spec, the
>>> > >> underlying problem is not solved. Without a good alternative or
>>> tooling
>>> > >> support, i am afraid this could become an adoption issue for v4
>>> going
>>> > >> forward.
>>> > >> >>
>>> > >> >> On 2026/07/17 17:13:55 huaxin gao wrote:
>>> > >> >> > Thanks all for the discussion. I've thought this over, and I'd
>>> like
>>> > >> to
>>> > >> >> > change my view to forbidding equality-delete writes for V4
>>> tables,
>>> > >> rather
>>> > >> >> > than the softer "deprecated but permitted."
>>> > >> >> >
>>> > >> >> > My earlier hesitation was about gating V4 on a write-time DV
>>> > >> implementation
>>> > >> >> > that isn't built yet. I think that concern goes away once we
>>> > >> separate the
>>> > >> >> > spec decision from engine adoption:
>>> > >> >> >
>>> > >> >> > - Forbidding equality deletes is a V4 format decision. It
>>> defines
>>> > >> what a V4
>>> > >> >> > table allows; it does not require every engine to have the
>>> > >> write-time DV
>>> > >> >> > path on day one.
>>> > >> >> > - Workloads that still rely on streaming upserts can stay on V3
>>> > >> until their
>>> > >> >> > engine's write-time path is ready, then adopt V4. Readers
>>> continue to
>>> > >> >> > support equality deletes for existing V2 and V3 tables.
>>> > >> >> > - So we don't need the write-time implementation finished to
>>> forbid
>>> > >> >> > equality deletes in V4. What we need is a clear, credible
>>> path, and
>>> > >> I think
>>> > >> >> > we have it: Flink's ConvertEqualityDeletes already maintains a
>>> PK
>>> > >> index in
>>> > >> >> > Flink state and does key -> position -> DV today, and the next
>>> step
>>> > >> is to
>>> > >> >> > persist that index into Iceberg so any engine can resolve
>>> positions
>>> > >> and
>>> > >> >> > emit DVs directly at write time.
>>> > >> >> >
>>> > >> >> > This also addresses Max's concern that "deprecated but
>>> permitted" is
>>> > >> too
>>> > >> >> > soft. Forbidding for V4 gives engines a real incentive to
>>> move, while
>>> > >> >> > keeping existing tables fully readable.
>>> > >> >> >
>>> > >> >> > On Xin's point, agreed that some perf data comparing DVs on V4
>>> > >> against
>>> > >> >> > equality deletes on V3 would be useful to have as we go.
>>> > >> >> >
>>> > >> >> > Thanks,
>>> > >> >> > Huaxin
>>> > >> >> >
>>> > >> >> > On Thu, Jul 16, 2026 at 8:43 AM Xin Huang via dev <
>>> > >> [email protected]>
>>> > >> >> > wrote:
>>> > >> >> >
>>> > >> >> > > Conceptually +1 as equality deletes really complicates the
>>> format
>>> > >> and
>>> > >> >> > > implementation.
>>> > >> >> > >
>>> > >> >> > > However given the concern is around performance side. Is
>>> there a
>>> > >> way to
>>> > >> >> > > make the decision making more data driven — having some
>>> > >> benchmarking on
>>> > >> >> > > perf comparison between dv+single file commit in v4 vs
>>> equality
>>> > >> delete in
>>> > >> >> > > v3 could help making a call.
>>> > >> >> > >
>>> > >> >> > > Thanks
>>> > >> >> > > Xin
>>> > >> >> > >
>>> > >> >> > > On Wed, Jul 15, 2026 at 11:01 PM Maximilian Michels <
>>> > >> [email protected]>
>>> > >> >> > > wrote:
>>> > >> >> > >
>>> > >> >> > >> I understood "deprecate equality deletes" as not forbidding
>>> > >> engines to
>>> > >> >> > >> write them, but rather discouraging them. IMHO this is long
>>> > >> overdue,
>>> > >> >> > >> but it is also a very soft transition. Perhaps too soft,
>>> because
>>> > >> it
>>> > >> >> > >> doesn't give engines who write them a real incentive to stop
>>> > >> writing
>>> > >> >> > >> equality deletes.
>>> > >> >> > >>
>>> > >> >> > >> Engines will likely be quicker to move away from writing
>>> equality
>>> > >> >> > >> deletes if we disallow writing them in V4. Regardless, we
>>> will
>>> > >> have to
>>> > >> >> > >> support reading equality deletes for V2 and V3 tables.
>>> > >> >> > >>
>>> > >> >> > >> I'm leaning more towards removing equality deletes for V4
>>> tables,
>>> > >> but
>>> > >> >> > >> I would like to hear what others think.
>>> > >> >> > >>
>>> > >> >> > >> -Max
>>> > >> >> > >>
>>> > >> >> > >>
>>> > >> >> > >> On Wed, Jul 15, 2026 at 10:03 PM Steven Wu <
>>> [email protected]>
>>> > >> wrote:
>>> > >> >> > >> >
>>> > >> >> > >> > > So I'd separate two things: deprecating in V4 (signal
>>> and
>>> > >> direction,
>>> > >> >> > >> safe to do now) versus forbidding equality-delete writes
>>> (gated
>>> > >> on the
>>> > >> >> > >> engine-agnostic path being ready). I'm only proposing the
>>> first
>>> > >> for V4.
>>> > >> >> > >> >
>>> > >> >> > >> > I thought we wanted to forbid equality-delete writes for
>>> v4
>>> > >> tables,
>>> > >> >> > >> which would really simplify the v4 adaptive metadata tree
>>> along
>>> > >> with other
>>> > >> >> > >> benefits that Huaxin already outlined,
>>> > >> >> > >> >
>>> > >> >> > >> > > deprecation path in v4
>>> > >> >> > >> >
>>> > >> >> > >> > I heard the Kafka connector in the Iceberg repo doesn't
>>> produce
>>> > >> >> > >> equality deletes. We would need to migrate the Flink sink to
>>> > >> leverage the
>>> > >> >> > >> index to produce DVs only in v4.
>>> > >> >> > >> >
>>> > >> >> > >> > On Wed, Jul 15, 2026 at 12:13 PM huaxin gao <
>>> > >> [email protected]>
>>> > >> >> > >> wrote:
>>> > >> >> > >> >>
>>> > >> >> > >> >> Thanks Max and Manu.
>>> > >> >> > >> >>
>>> > >> >> > >> >> Max, thanks for the added detail. It's a good point that
>>> the
>>> > >> index in
>>> > >> >> > >> ConvertEqualityDeletes is persisted in Flink state
>>> (RocksDB) and
>>> > >> updated
>>> > >> >> > >> incrementally. That strengthens the case, since it shows
>>> the key
>>> > >> to
>>> > >> >> > >> position index is already durable, just scoped to Flink
>>> today.
>>> > >> And your
>>> > >> >> > >> closing point is exactly the plan I have in mind: deprecate
>>> > >> equality
>>> > >> >> > >> deletes in V4, and once the spec has an index, persist that
>>> PK
>>> > >> index into
>>> > >> >> > >> Iceberg so it can be shared across engines.
>>> > >> >> > >> >>
>>> > >> >> > >> >> Manu, good question on how this works in practice. To be
>>> > >> clear, I'm
>>> > >> >> > >> proposing "deprecated but permitted," not removal. In V4,
>>> writers
>>> > >> >> > >> (including Kafka Connect) could keep emitting equality
>>> deletes
>>> > >> and readers
>>> > >> >> > >> would keep applying them. Deprecation just declares deletion
>>> > >> vectors the
>>> > >> >> > >> going-forward mechanism and stops new investment in equality
>>> > >> deletes. So V4
>>> > >> >> > >> would not be gated on the index landing.
>>> > >> >> > >> >>
>>> > >> >> > >> >> On the cleanup path today: ConvertEqualityDeletes runs
>>> against
>>> > >> the
>>> > >> >> > >> table, not a specific writer, so a Kafka Connect pipeline
>>> can
>>> > >> already be
>>> > >> >> > >> cleaned up. It keeps writing equality deletes, and a
>>> standalone
>>> > >> Flink
>>> > >> >> > >> ConvertEqualityDeletes job converts them to DVs. The one
>>> friction
>>> > >> is that
>>> > >> >> > >> the conversion runtime is Flink today, so a Kafka-only shop
>>> would
>>> > >> have to
>>> > >> >> > >> run Flink just for maintenance. A Spark action would be a
>>> natural
>>> > >> follow-up
>>> > >> >> > >> here, since Spark is the usual Iceberg maintenance engine
>>> and
>>> > >> most batch
>>> > >> >> > >> shops already run it.
>>> > >> >> > >> >>
>>> > >> >> > >> >> Longer term, the next step is to persist the key to
>>> position
>>> > >> index
>>> > >> >> > >> into Iceberg. Once it's shared, engines can look up
>>> positions and
>>> > >> write DVs
>>> > >> >> > >> directly at write time, so a writer can stop producing
>>> equality
>>> > >> deletes
>>> > >> >> > >> entirely, and neither the Flink job nor a Spark action
>>> needs to
>>> > >> rebuild the
>>> > >> >> > >> index each run. Each engine (Kafka Connect, Spark, Flink)
>>> would
>>> > >> adopt
>>> > >> >> > >> write-time DVs on its own schedule; until then it keeps
>>> writing
>>> > >> equality
>>> > >> >> > >> deletes and relies on background conversion. So V4
>>> deprecation
>>> > >> doesn't
>>> > >> >> > >> require re-implementing every writer up front.
>>> > >> >> > >> >>
>>> > >> >> > >> >> So I'd separate two things: deprecating in V4 (signal and
>>> > >> direction,
>>> > >> >> > >> safe to do now) versus forbidding equality-delete writes
>>> (gated
>>> > >> on the
>>> > >> >> > >> engine-agnostic path being ready). I'm only proposing the
>>> first
>>> > >> for V4.
>>> > >> >> > >> >>
>>> > >> >> > >> >> Thanks,
>>> > >> >> > >> >> Huaxin
>>> > >> >> > >> >>
>>> > >> >> > >> >> On Wed, Jul 15, 2026 at 3:10 AM Manu Zhang <
>>> > >> [email protected]>
>>> > >> >> > >> wrote:
>>> > >> >> > >> >>>
>>> > >> >> > >> >>> Hi Huaxin,
>>> > >> >> > >> >>>
>>> > >> >> > >> >>> +1 for deprecating equality deletes, but how would this
>>> > >> deprecation
>>> > >> >> > >> work practically in V4?
>>> > >> >> > >> >>> As Max pointed out, we still lack an engine-agnostic
>>> solution
>>> > >> for
>>> > >> >> > >> streaming use cases. For example, how would we handle
>>> equality
>>> > >> deletes
>>> > >> >> > >> written by Kafka Connect?
>>> > >> >> > >> >>> While the index proposal looks promising, I don't see a
>>> clear
>>> > >> path
>>> > >> >> > >> for deprecating equality deletes in V4 before that index
>>> work
>>> > >> actually
>>> > >> >> > >> lands.
>>> > >> >> > >> >>>
>>> > >> >> > >> >>> Thanks,
>>> > >> >> > >> >>> Manu
>>> > >> >> > >> >>>
>>> > >> >> > >> >>>
>>> > >> >> > >> >>> On Wed, Jul 15, 2026 at 5:30 PM Maximilian Michels <
>>> > >> [email protected]>
>>> > >> >> > >> wrote:
>>> > >> >> > >> >>>>
>>> > >> >> > >> >>>> Hi Huaxin,
>>> > >> >> > >> >>>>
>>> > >> >> > >> >>>> Thanks for reviving the discussion on deprecating
>>> equality
>>> > >> deletes.
>>> > >> >> > >> >>>> Equality deletes are the number one pain for streaming
>>> use
>>> > >> cases.
>>> > >> >> > >> Many
>>> > >> >> > >> >>>> users give up when they see the merge-on-read costs,
>>> or they
>>> > >> build
>>> > >> >> > >> >>>> custom solutions which move them further away from core
>>> > >> Iceberg. That
>>> > >> >> > >> >>>> said, we've made great progress since the initial
>>> > >> conversation in
>>> > >> >> > >> >>>> 2024.
>>> > >> >> > >> >>>>
>>> > >> >> > >> >>>> Just to add what you said: The index we maintain in
>>> > >> >> > >> >>>> ConvertEqualityDeletes is not ephemeral. The index is
>>> > >> persisted in
>>> > >> >> > >> >>>> Flink's managed state (RocksDB). It is continuously
>>> updated
>>> > >> as new
>>> > >> >> > >> >>>> data arrives and checkpointed periodically. However,
>>> even
>>> > >> though the
>>> > >> >> > >> >>>> conversion works for data written by any engine, we
>>> > >> currently require
>>> > >> >> > >> >>>> Flink for the conversion itself. Storing the index
>>> directly
>>> > >> in
>>> > >> >> > >> Iceberg
>>> > >> >> > >> >>>> and enabling all engines access would be the next
>>> logical
>>> > >> step
>>> > >> >> > >> towards
>>> > >> >> > >> >>>> a fully engine-agnostic solution.
>>> > >> >> > >> >>>>
>>> > >> >> > >> >>>> The reality is that we don't yet have a working
>>> solution to
>>> > >> avoid
>>> > >> >> > >> >>>> writing equality deletes across all engines, but given
>>> the
>>> > >> recent
>>> > >> >> > >> >>>> progress, the proposed plan seems realistic. So +1 for
>>> > >> deprecating
>>> > >> >> > >> >>>> equality deletes in V4.
>>> > >> >> > >> >>>>
>>> > >> >> > >> >>>> Cheers,
>>> > >> >> > >> >>>> Max
>>> > >> >> > >> >>>>
>>> > >> >> > >> >>>>
>>> > >> >> > >> >>>>
>>> > >> >> > >> >>>> On Tue, Jul 14, 2026 at 3:25 AM huaxin gao <
>>> > >> [email protected]>
>>> > >> >> > >> wrote:
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > Hi all,
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > I'd like to restart the conversation about
>>> deprecating
>>> > >> equality
>>> > >> >> > >> deletes, now in the context of the V4 spec.
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > Background
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > This isn't a new idea. Russell proposed deprecating
>>> > >> equality
>>> > >> >> > >> deletes in V3 and removing them from the spec in V4, back in
>>> > >> October 2024
>>> > >> >> > >> in "[DISCUSS] - Deprecate Equality Deletes". The main
>>> blocker at
>>> > >> the time
>>> > >> >> > >> was that equality deletes served real use cases (especially
>>> Flink
>>> > >> streaming
>>> > >> >> > >> upserts) with no efficient alternative. Two developments
>>> since
>>> > >> then make
>>> > >> >> > >> the V4 removal worth acting on now.
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > Why equality deletes are costly
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > Equality deletes are cheap to write but expensive to
>>> read:
>>> > >> a
>>> > >> >> > >> reader must load the equality-delete files and join them
>>> against
>>> > >> every
>>> > >> >> > >> candidate row in the delete's sequence-number range.
>>> Positional
>>> > >> deletes
>>> > >> >> > >> skip that per-row join by marking exact positions, so they
>>> have
>>> > >> always read
>>> > >> >> > >> faster, and V3 deletion vectors make them faster still, one
>>> > >> compact bitmap
>>> > >> >> > >> per data file, applied by an O(1) position check, instead
>>> of V2's
>>> > >> many
>>> > >> >> > >> position-delete files. So equality deletes' only real edge
>>> is the
>>> > >> cheap
>>> > >> >> > >> write, and both background conversion and a write-time
>>> key-lookup
>>> > >> index can
>>> > >> >> > >> recover that.
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > Beyond performance
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > Equality deletes also block other features. CDC and
>>> row
>>> > >> lineage
>>> > >> >> > >> are effectively impossible while they are in use, because
>>> the
>>> > >> true state of
>>> > >> >> > >> the table can only be determined with a full scan. That same
>>> > >> property means
>>> > >> >> > >> differential structures such as materialized views and
>>> secondary
>>> > >> indexes
>>> > >> >> > >> have to be fully rebuilt whenever an equality delete is
>>> added,
>>> > >> rather than
>>> > >> >> > >> maintained incrementally. So removing equality deletes is
>>> close
>>> > >> to a
>>> > >> >> > >> prerequisite for the index work to stay incrementally
>>> > >> maintainable.
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > Evidence the alternatives are practical
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > 1. Converting equality deletes to DVs works today.
>>> Max
>>> > >> Michels'
>>> > >> >> > >> ConvertEqualityDeletes maintenance task (16831, 16844,
>>> 16858,
>>> > >> 16874, 16889,
>>> > >> >> > >> 16948) rewrites equality deletes into deletion vectors as a
>>> > >> background
>>> > >> >> > >> Flink job: the writer keeps appending equality deletes to a
>>> > >> staging branch,
>>> > >> >> > >> and the task converts them to DVs on the target branch so
>>> reads
>>> > >> apply
>>> > >> >> > >> deletes by position. Notably, the task resolves each delete
>>> to a
>>> > >> position
>>> > >> >> > >> using a primary-key index that it builds and maintains
>>> inside the
>>> > >> job,
>>> > >> >> > >> demonstrating the full "key -> position -> DV" path end to
>>> end.
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > 2. A persistent key-lookup index removes the need to
>>> write
>>> > >> them at
>>> > >> >> > >> all. The secondary index spec we're working on (#16961)
>>> includes a
>>> > >> >> > >> key-lookup index mapping a key to its data file and row
>>> position.
>>> > >> This is
>>> > >> >> > >> essentially the persistent, catalog-managed form of the
>>> index
>>> > >> Max's task
>>> > >> >> > >> builds ephemerally. With it, a writer can resolve positions
>>> at
>>> > >> write time
>>> > >> >> > >> and emit DVs directly, without ever producing an equality
>>> delete.
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > How these two efforts fit together
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > They're complementary, and they cover the two things
>>> we
>>> > >> need to
>>> > >> >> > >> deprecate equality deletes:
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > Migration (existing data): ConvertEqualityDeletes
>>> cleans
>>> > >> up tables
>>> > >> >> > >> that already contain equality deletes, and supports writers
>>> that
>>> > >> still emit
>>> > >> >> > >> them, converting them to DVs in the background.
>>> > >> >> > >> >>>> > Going forward (new writes): the persistent key-lookup
>>> > >> index lets
>>> > >> >> > >> writers skip equality deletes entirely by looking up
>>> positions
>>> > >> directly.
>>> > >> >> > >> >>>> > The connection is that Max's task already proves the
>>> core
>>> > >> >> > >> mechanism (resolve key -> position, write a DV); it just
>>> rebuilds
>>> > >> a
>>> > >> >> > >> throwaway index each cycle. A durable, shared index both
>>> enables
>>> > >> write-time
>>> > >> >> > >> elimination and removes that rebuild cost from the
>>> conversion
>>> > >> path.
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > Proposal
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > I propose that we deprecate equality deletes in V4.
>>> The
>>> > >> blocker
>>> > >> >> > >> from 2024 was the lack of a viable alternative, and we now
>>> have
>>> > >> the pieces:
>>> > >> >> > >> background conversion to DVs works today, and the key-lookup
>>> > >> index gives us
>>> > >> >> > >> a path to eliminating them at write time. Deletion vectors
>>> should
>>> > >> be the
>>> > >> >> > >> going-forward mechanism for row-level deletes and upserts,
>>> > >> produced by
>>> > >> >> > >> background conversion now and directly by writers once the
>>> index
>>> > >> is
>>> > >> >> > >> available. Readers would continue to support equality
>>> deletes for
>>> > >> backward
>>> > >> >> > >> compatibility with existing V2/V3 tables.
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > Migration path
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > Existing tables keep working; readers continue to
>>> apply
>>> > >> equality
>>> > >> >> > >> deletes.
>>> > >> >> > >> >>>> > ConvertEqualityDeletes (Flink) rewrites existing
>>> equality
>>> > >> deletes
>>> > >> >> > >> into DVs so tables can be cleared of them over time.
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > I'd love people's thoughts, especially from those
>>> running
>>> > >> large
>>> > >> >> > >> streaming-upsert workloads.
>>> > >> >> > >> >>>> >
>>> > >> >> > >> >>>> > Thanks,
>>> > >> >> > >> >>>> > Huaxin
>>> > >> >> > >>
>>> > >> >> > >
>>> > >> >> >
>>> > >>
>>> > >>
>>> > >>
>>> > >> --
>>> > >> Regards
>>> > >> Junwang Zhao
>>> > >>
>>> > >
>>> >
>>>
>>

Reply via email to