Hi Neelesh,

Thanks for the suggestion. I agree that a sync would be helpful.

If I remember correctly, when the proposal was brought up in the Iceberg 
Community Sync on 8/5 there was general support for the idea of creating a 
shared repository for testing/verification/conformance fixture.

There were a few smaller open questions around repository naming, and whether 
we should include a mechanism to verify the correctness of written artifacts 
through the new repository as well.

We can go through these items and other questions in the sync, and hopefully 
kick of the infrastructure process to set up the repository.

Sung

On 2026/08/13 21:25:05 Neelesh Salian wrote:
> Hi folks,
> 
> To help gather folks to talk about the proposal [1] in sync, I've set up a
> sync on 8/19/2026 (Wednesday) at 9 am PT.
> Here is the invite link: https://calendar.app.google/cpos7A4jf1Ko9C6z8. It
> should be on the Iceberg Dev Events Calendar as well.
> The session will be recorded.
> 
> Thank you.
> 
> [1] Shared test fixtures for Apache Iceberg implementations Proposal
> <https://docs.google.com/document/d/1Qr82hccxJhWxKRT7B-3ylhOLSaJTYfc7hU51-mhzK6Y/edit?usp=sharing>
> 
> On Mon, Aug 10, 2026 at 2:21 PM Andrei Tserakhau via dev <
> [email protected]> wrote:
> 
> > Hi all,
> >
> > +1 on the read-first corpus - it's the right v1 and I'm glad we converged
> > on it. I would vote for any generic name that allows us to pursue
> > long-term goals.
> > vision of DMLs; my personal favorite is iceberg-verification.
> >
> > One thing I want to put on the table, tied to the read-vs-write
> > question from the sync.
> >
> > A lot of the sharpest bugs in our own motivation list are write bugs, not
> > read bugs. equality_ids typed as long instead of int (#880), or a
> > merge-on-read delete that collapses to a plain file removal in one
> > implementation where another writes a delete file - those live in what an
> > implementation *produces*, and a read-only fixture (which ships a
> > known-good file) never exercises the code that emits them. So write
> > verification isn't only a "later" item; it's where a good share of the
> > divergence actually lives.
> >
> > The prototype referenced in the doc [3] is an exploration of exactly that,
> > and it's engine-free. Describe DML as a logical, engine-agnostic op-log
> > (append / delete / evolve-schema / ...), have each implementation execute
> > it through its own API and emit a canonical decoded form, and compare that
> > centrally against the same field-id-keyed expected value the read fixtures
> > use. A read fixture is just this frozen to disk, so one comparator serves
> > both directions.
> >
> > There's a live demo of where it goes:
> >
> > https://laskoviymishka.github.io/iceberg-verification/
> >
> > A cross-implementation matrix over go, rust and java, covering v1-v3 on
> > both read and write. No Spark in the loop - the Java reference mints each
> > case and the readers diff against it (this is the engine-free orchestrator
> > + matrix the alternatives section calls feasible-but-deferred; the demo is
> > the evidence it's feasible, and it builds three client runners, not five
> > toolchains). It also runs a small differential fuzz campaign: generate
> > random valid op-logs, mint each with the Java reference, diff every reader.
> > Even a bounded run already surfaced some bugs - one reader rendered
> > a fixed type as fixed(4) where the Java reference produced fixed[4], which
> > is precisely the type-string surface we sequenced first. And an int-to-long
> > promotion was implemented in a runner working only from the fixtures - the
> > fixtures-as-spec loop working end to end.
> >
> > Where I think this eventually points, and the reason the op-log is logical
> > rather than byte-level: the same fixtures don't have to stop at language
> > libraries. Because the ops are engine-agnostic, an engine (Spark, Trino,
> > ...) can execute the same op-log and emit the same canonical form, so
> > conformance can cover engines too, not just clients. Combined with
> > spec-first development - a new spec feature ships with the fixtures that
> > define it - a library and an engine both prove they conform by making the
> > same fixtures pass. That's the long-term shape and north-star we can
> > aim after V1 landed.
> >
> > Happy to walk through the demo and the op-log format in a sync if there's
> > interest.
> >
> > Thanks,
> > Andrei
> >
> > On Mon, Aug 10, 2026 at 1:25 AM Neelesh Salian <[email protected]>
> > wrote:
> >
> >> Hi all,
> >>
> >> Thank you for all the discussion this week on the community sync.
> >> We had a couple of items raised around:
> >> - The repository naming
> >> - Verifying the implementations' correctness via the artifacts on both
> >> reads and writes.
> >>
> >> Please take a look and add your feedback on the proposal [1] document
> >> for any suggestions, thoughts, or enhancements.
> >> In addition, I'll check back in a week and see where the proposal goes,
> >> and maybe set up a dedicated sync if folks would like to dive deeper on
> >> this.
> >>
> >> Thanks,
> >> Neelesh, Sung, Andrei
> >>
> >>
> >> [1] Shared test fixtures for Apache Iceberg implementations Proposal
> >> <https://docs.google.com/document/d/1Qr82hccxJhWxKRT7B-3ylhOLSaJTYfc7hU51-mhzK6Y/edit?usp=sharing>
> >>
> >>
> >>
> >> On Tue, Jul 14, 2026 at 7:16 AM Neelesh Salian <[email protected]>
> >> wrote:
> >>
> >>> Hi all,
> >>>
> >>> Sung and I each started a thread on this recently [1][2], and together
> >>> with Andrei the discussion converged on a single design.
> >>>
> >>> We've merged them into one proposal: a standalone, language-neutral
> >>> repository of conformance fixtures - starting with static inputs paired
> >>> with the expected value the spec fixes - modeled on parquet-testing, so
> >>> each implementation checks its reading of the spec against a shared set
> >>> instead of only against itself.
> >>>
> >>> There are working POCs: Sung's iceberg-testing fork [3] with pyiceberg
> >>> [4] and
> >>> iceberg-rust [5], and iceberg-go's engine-free variant checks [6].
> >>>
> >>> Document: [Proposal
> >>> <https://docs.google.com/document/d/1Qr82hccxJhWxKRT7B-3ylhOLSaJTYfc7hU51-mhzK6Y/edit?usp=sharing>
> >>> ]
> >>>
> >>> The plan: start with the smallest primitives (type strings, transforms),
> >>> grow outward
> >>> to binary fixtures; no harness and no central gate - each implementation
> >>> pins the repo
> >>> and asserts in its own suite; Java and Go first, covering v2 and v3.
> >>>
> >>> Feedback welcome on this and happy to set up a short sync for anyone who
> >>> wants to dig in.
> >>>
> >>> Thanks,
> >>> Neelesh, Sung, Andrei
> >>>
> >>> [1] https://lists.apache.org/thread/964630c6q0jovs579x1jzb1t0o19jgjg
> >>> [2] https://lists.apache.org/thread/87cvl9gk0cjk1of7jh3nvm4lzvzxnc8m
> >>> [3] https://github.com/sungwy/iceberg-testing
> >>> [4] https://github.com/sungwy/iceberg-python/pull/1
> >>> [5] https://github.com/sungwy/iceberg-rust/pull/2
> >>> [6] https://github.com/apache/iceberg-go/pull/1238
> >>>
> >>>
> 

Reply via email to