Hi folks,

To help gather folks to talk about the proposal [1] in sync, I've set up a
sync on 8/19/2026 (Wednesday) at 9 am PT.
Here is the invite link: https://calendar.app.google/cpos7A4jf1Ko9C6z8. It
should be on the Iceberg Dev Events Calendar as well.
The session will be recorded.

Thank you.

[1] Shared test fixtures for Apache Iceberg implementations Proposal
<https://docs.google.com/document/d/1Qr82hccxJhWxKRT7B-3ylhOLSaJTYfc7hU51-mhzK6Y/edit?usp=sharing>

On Mon, Aug 10, 2026 at 2:21 PM Andrei Tserakhau via dev <
[email protected]> wrote:

> Hi all,
>
> +1 on the read-first corpus - it's the right v1 and I'm glad we converged
> on it. I would vote for any generic name that allows us to pursue
> long-term goals.
> vision of DMLs; my personal favorite is iceberg-verification.
>
> One thing I want to put on the table, tied to the read-vs-write
> question from the sync.
>
> A lot of the sharpest bugs in our own motivation list are write bugs, not
> read bugs. equality_ids typed as long instead of int (#880), or a
> merge-on-read delete that collapses to a plain file removal in one
> implementation where another writes a delete file - those live in what an
> implementation *produces*, and a read-only fixture (which ships a
> known-good file) never exercises the code that emits them. So write
> verification isn't only a "later" item; it's where a good share of the
> divergence actually lives.
>
> The prototype referenced in the doc [3] is an exploration of exactly that,
> and it's engine-free. Describe DML as a logical, engine-agnostic op-log
> (append / delete / evolve-schema / ...), have each implementation execute
> it through its own API and emit a canonical decoded form, and compare that
> centrally against the same field-id-keyed expected value the read fixtures
> use. A read fixture is just this frozen to disk, so one comparator serves
> both directions.
>
> There's a live demo of where it goes:
>
> https://laskoviymishka.github.io/iceberg-verification/
>
> A cross-implementation matrix over go, rust and java, covering v1-v3 on
> both read and write. No Spark in the loop - the Java reference mints each
> case and the readers diff against it (this is the engine-free orchestrator
> + matrix the alternatives section calls feasible-but-deferred; the demo is
> the evidence it's feasible, and it builds three client runners, not five
> toolchains). It also runs a small differential fuzz campaign: generate
> random valid op-logs, mint each with the Java reference, diff every reader.
> Even a bounded run already surfaced some bugs - one reader rendered
> a fixed type as fixed(4) where the Java reference produced fixed[4], which
> is precisely the type-string surface we sequenced first. And an int-to-long
> promotion was implemented in a runner working only from the fixtures - the
> fixtures-as-spec loop working end to end.
>
> Where I think this eventually points, and the reason the op-log is logical
> rather than byte-level: the same fixtures don't have to stop at language
> libraries. Because the ops are engine-agnostic, an engine (Spark, Trino,
> ...) can execute the same op-log and emit the same canonical form, so
> conformance can cover engines too, not just clients. Combined with
> spec-first development - a new spec feature ships with the fixtures that
> define it - a library and an engine both prove they conform by making the
> same fixtures pass. That's the long-term shape and north-star we can
> aim after V1 landed.
>
> Happy to walk through the demo and the op-log format in a sync if there's
> interest.
>
> Thanks,
> Andrei
>
> On Mon, Aug 10, 2026 at 1:25 AM Neelesh Salian <[email protected]>
> wrote:
>
>> Hi all,
>>
>> Thank you for all the discussion this week on the community sync.
>> We had a couple of items raised around:
>> - The repository naming
>> - Verifying the implementations' correctness via the artifacts on both
>> reads and writes.
>>
>> Please take a look and add your feedback on the proposal [1] document
>> for any suggestions, thoughts, or enhancements.
>> In addition, I'll check back in a week and see where the proposal goes,
>> and maybe set up a dedicated sync if folks would like to dive deeper on
>> this.
>>
>> Thanks,
>> Neelesh, Sung, Andrei
>>
>>
>> [1] Shared test fixtures for Apache Iceberg implementations Proposal
>> <https://docs.google.com/document/d/1Qr82hccxJhWxKRT7B-3ylhOLSaJTYfc7hU51-mhzK6Y/edit?usp=sharing>
>>
>>
>>
>> On Tue, Jul 14, 2026 at 7:16 AM Neelesh Salian <[email protected]>
>> wrote:
>>
>>> Hi all,
>>>
>>> Sung and I each started a thread on this recently [1][2], and together
>>> with Andrei the discussion converged on a single design.
>>>
>>> We've merged them into one proposal: a standalone, language-neutral
>>> repository of conformance fixtures - starting with static inputs paired
>>> with the expected value the spec fixes - modeled on parquet-testing, so
>>> each implementation checks its reading of the spec against a shared set
>>> instead of only against itself.
>>>
>>> There are working POCs: Sung's iceberg-testing fork [3] with pyiceberg
>>> [4] and
>>> iceberg-rust [5], and iceberg-go's engine-free variant checks [6].
>>>
>>> Document: [Proposal
>>> <https://docs.google.com/document/d/1Qr82hccxJhWxKRT7B-3ylhOLSaJTYfc7hU51-mhzK6Y/edit?usp=sharing>
>>> ]
>>>
>>> The plan: start with the smallest primitives (type strings, transforms),
>>> grow outward
>>> to binary fixtures; no harness and no central gate - each implementation
>>> pins the repo
>>> and asserts in its own suite; Java and Go first, covering v2 and v3.
>>>
>>> Feedback welcome on this and happy to set up a short sync for anyone who
>>> wants to dig in.
>>>
>>> Thanks,
>>> Neelesh, Sung, Andrei
>>>
>>> [1] https://lists.apache.org/thread/964630c6q0jovs579x1jzb1t0o19jgjg
>>> [2] https://lists.apache.org/thread/87cvl9gk0cjk1of7jh3nvm4lzvzxnc8m
>>> [3] https://github.com/sungwy/iceberg-testing
>>> [4] https://github.com/sungwy/iceberg-python/pull/1
>>> [5] https://github.com/sungwy/iceberg-rust/pull/2
>>> [6] https://github.com/apache/iceberg-go/pull/1238
>>>
>>>

Reply via email to