I strongly support removing Equality deletes in V4 (although this is probably very obvious from my previous thread trying to remove them in V3 :)
I actually don't think we even need to ensure there are replacements in V4 at ratification time as long as we have a path forward for implementation. Too many V4 benefits are essentially completely removed if someone writes an equality delete and I think we are much better off without them. I wouldn't require that a table with a V4 root manifest to have *no* equality deletes (equality deletes in a V3 Delete manifest are fine), but there should be no mechanism for adding or moving an equality delete directly to V4 Metadata. On Fri, Jul 24, 2026 at 3:14 AM Péter Váry <[email protected]> wrote: > I fully support the direction of disallowing new equality deletes in V4. > Max's work on the equality delete resolver demonstrates that a viable > replacement approach is possible. > > That said, there are a few prerequisites I would like to see addressed > before we completely move away from equality deletes: > > - The current solution still writes equality deletes to a staging > branch. We need an approach that avoids committing equality deletes to the > table altogether. > - The index needs to be written directly from the Flink job and reused > across job restarts. This is important both for fast job initialization and > efficient cache warmup. We are still actively working on the index > specification required to support this. > > There are also a couple of areas that are not blockers but would be > valuable improvements: > > - Support for row IDs in streaming jobs. The current solution does not > handle this, and solving it is non-trivial. > - Incremental cache refresh on concurrent commits. Today, the > implementation performs a full index refresh, but there is room for a more > efficient approach. > > Since both the V4 specification and implementation are still in progress, > I think we are in a good position to proceed. However, we should ensure > these pieces come together along the way so that by the time V4 is > released, we already have a complete and production-ready replacement for > equality deletes. > > Thanks, > Peter > > vaquar khan <[email protected]> ezt írta (időpont: 2026. júl. 24., P, > 9:05): > >> Hi all, >> Thanks Huaxin for restarting this, and Max for the detail on the Flink >> index. A few people asked for numbers comparing deletion vectors against >> equality deletes (Xin, Huaxin), and Xiening asked whether this just moves >> the cost from the reader to the writer. I already had these benchmarks set >> up, so I ran Iceberg's own JMH benchmarks, unchanged, on AWS EC2. Full >> data, charts, and setup are here: >> >> https://docs.google.com/document/d/12ifnKwiWemXcLndv19QJXOxh5o4lOsRxw_HsOgQ4sdU/edit?tab=t.0 >> Short version (Iceberg's own read-path and DV-reader benchmarks, >> unmodified, on an r7i.4xlarge; Iceberg commit 2f17d9b, OpenJDK 17.0.19, JMH >> SingleShotTime, 3 warm-up + 5 measured iterations): >> >> - Read cost grows as more rows are deleted. For a 10,000,000-row read >> (readIcebergVectorized, s/op), equality vs position/DV is about 1.0x at 0 >> percent deleted, 1.27x at 5 percent, 1.38x at 25 percent, and about 2.6x >> at >> 100 percent. The 100 percent point is noisy, so I read it as a direction, >> not a precise figure; the 5 to 25 percent points and the upward trend are >> the solid part. >> - Applying the delete is flat for deletion vectors and not flat for >> position-delete files. In the DV reader benchmark, DV load stays around >> 0.0007 to 0.0009 s/op regardless of delete density, while partition-scoped >> position deletes grow to about 660x that at 20 percent deleted(0.519 vs >> 0.00079 s/op). >> >> *On Xiening's point*: the reader cost is real, and I measured it. With >> equality deletes, every reader pays the join on every query, for as long as >> the table lives. Resolving the position once (at write time, or once in a >> background conversion) turns every later read into the flat deletion-vector >> check. So to me it reads less like "move the burden to the writer" and more >> like "pay it once instead of on every read," which is the same shape as the >> v3 move from many position-delete files to one deletion vector. >> *What I did not measure, to be clear:* the cost of building and keeping >> the write-time index, or the streaming write path. Max already covered that >> side with the ConvertEqualityDeletes RocksDB primary-key index. I have >> short answers to Junwang's two index questions in the doc as well. >> *Where I land:* +1 to deprecating equality deletes in V4, and I lean >> toward the stronger version several of you have proposed, forbidding new >> equality-delete writes in V4 tables while keeping them readable for V2 and >> V3. My reason is the read-path data above: the extra cost of equality >> deletes is structural and is paid on every read, and a deletion vector >> turns that into a flat, one-time cost, so making DVs the going-forward >> mechanism is the right call. >> *One honest limit on my +1:* my numbers are about the read side. They do >> not measure the write-time index cost that Xiening raised, so I am not >> claiming the write path is free. That is why I would pair forbidding in V4 >> with the migration path already discussed here: streaming-upsert workloads >> stay on V3 (writing equality deletes, with background conversion keeping >> reads fast) until their engine can emit DVs directly, and non-streaming >> workloads move to V4 right away. Max's Flink RocksDB index suggests the >> write-time path is tractable. >> *One side note from running these:* on JDK 17 the JMH harness does not >> run cleanly. The Spark read benchmarks fail at warm-up without a few extra >> module opens (sun.util.calendar and friends), and the full-scale planning >> benchmarks run out of memory at the built-in 32g heap. This is a >> harness/tooling issue, not an Iceberg correctness problem, and I filed it >> with a proposed fix: >> https://github.com/apache/iceberg/issues/17330 >> Happy to share the raw JMH result JSON if anyone wants it, and glad to >> re-run any specific case (delete density, row count, column mix) that would >> help. >> Regards, >> Viquar Khan >> https://www.linkedin.com/in/vaquar-khan-b695577/ >> >> >> On Thu, 23 Jul 2026 at 19:50, Steve <[email protected]> wrote: >> >>> I want to +1 to forbid new equality-delete writes in v4, keep them >>> readable for backward compatibility. Equality delete imposes an asymetrics >>> burden to the reader while simplifying the writer's work, also slows down >>> and complicates the feature development. >>> >>> For existing v2/v3 tables with equality deletes, upgrading to v4 should >>> continue to be a lightweight, metadata-only operation, with no synchronous >>> EQ-delete rewrite required. We have a working precedent for exactly this >>> shape of change: how DVs were introduced and adopted in v3. A v2 table with >>> existing position-delete files could upgrade to v3 without rewriting >>> anything; those files stayed valid for reads. The spec requirement drove >>> the engine change instead: writers had to produce deletion vectors for new >>> DML going forward. As Huaxin proposed, converting the carried-over equality >>> deletes into DVs can be a separate maintenance action. >>> >>> Thanks, >>> Hongyue >>> >>
