andygrove opened a new pull request, #6417: URL: https://github.com/apache/datafusion-comet/pull/6417
## Which issue does this PR close? N/A. There is no tracking issue for the promotion. ## Rationale for this change Since #4950, Spark 4.2 has the same CI coverage as Spark 3.5 and 4.0: Comet's own suites and Apache Spark's SQL suite both run nightly against `main`. The docs still call 4.2 experimental and say it should not be used in production. The docs also hid a packaging gap. No Spark 4.2 jar has ever been published: neither `dev/release/build-release-comet.sh` nor `publish_snapshot.yml` builds the `spark-4.2` profile, and Maven Central has no `comet-spark-spark4.2_2.13` artifact. The Gluten comparison's "experimental build also published for the Spark 4.2 preview" was never true. Where 4.2 stands against the promotion criteria in Stage 6 of the new-version guide: - 4.2.0 is a final upstream release. - The Spark SQL suite has run nightly only since #4950 merged on 2026-09-28, which is short of the guide's one release cycle. The first nightly to include it (#6380) failed before running any test, because the 4.2 diff still named Comet `1.1.0-SNAPSHOT`. #6398 addresses that, so this PR should merge after #6398 lands and a nightly passes. - `assume(!isSpark42Plus)` skips remain for #4967 (ANSI overflow error tests, labelled `correctness`), #4968 (BloomFilter tests) and #4969 (Iceberg REST catalog, which is infrastructure). - #4966 (`collect_set` NaN and -0.0 normalization) is open with `priority:critical`. #5166 implemented the Spark 4.2 normalization. What remains is re-enabling the Spark test in the 4.2 diff, which #5209 tracks. ## What changes are included in this PR? User guide and README: - `installation.md`: 4.2.0 moves into the supported table (Java 17, Scala 2.13, Nightly / Nightly). The experimental table goes away because 4.2 was its only row. The snapshot artifact list and the release download links gain the Spark 4.2 jar. - `spark-versions.md`: the 4.2 section drops its experimental warning and gains a known limitations list: window aggregates with `FILTER (WHERE ...)` fall back, a `OneRowRelation` in `UNION` branches takes the union and the aggregates above it off Comet (#4949), and the ANSI overflow differences (#4967). - `README.md` and the Gluten comparison list 4.2 as supported. - A Spark 4.2 expression compatibility page: `compatibility/expressions/spark-4.2/index.md` plus its toctree entry. #6168 noted that it was missing. Packaging and doc generation: - `dev/release/build-release-comet.sh` and `publish_snapshot.yml` build and publish `comet-spark-spark4.2_2.13`, so 1.2.0 ships the jar. - `docs/build.sh` and `dev/generate-release-docs.sh` generate the 4.2 pages. `spark-4.2` goes before `spark-4.1`, not after it. Every profile run rewrites the shared `configs.md` and `expressions.md`, so the last profile wins. Spark 4.2 registers `max_by` and `min_by` through `MaxByBuilder` and `MinByBuilder`, which the generator cannot map to a serde, so with 4.2 last both rows would show the unknown placeholder instead of `Native`. Contributor docs: - The workflows README and the new-version guide no longer describe 4.2 as experimental. Stage 6 of the guide now lists the packaging and doc-generation steps above, so the next promotion does not miss them. Left alone: comments in `pom.xml`, `ci.yml` and `compute-changes.py` still call 4.2 experimental. Touching `pom.xml` or `ci.yml` runs the full merge-queue suite (macOS, Spark SQL 4.1, Iceberg 1.11 and more) for a comment change, so those comments can ride along with the next change to those files. As it stands, this PR runs no heavy CI job. This targets `main` only. `branch-1.1` does not have #4950, including the window `FILTER` fallback, so Comet 1.1.0 on Spark 4.2 would evaluate those aggregates without their filter. The change should not be backported. ## How are these changes tested? - Ran the docs generator for `spark-4.2` the way `docs/build.sh` does. All ten category pages generate. Compared with a `spark-4.1` run of the shared output: `configs.md` is identical, and `expressions.md` differs only in the `max_by` and `min_by` rows described above. - Built the docs with Sphinx on this branch and on `main`. The only new warnings are for the ten spark-4.2 category pages, which exist only after generation. The other versions' pages produce the same warnings in a source-only build. With the generated pages copied in, the build has no spark-4.2 warnings. - `prettier --check` (3.9.9, as in CI), `dev/ci/check-ci-config.py`, `bash -n` on the three scripts, and a YAML parse of `publish_snapshot.yml` all pass. `compute-changes.py` selects no heavy job for this change set on either `pull_request` or `merge_group`. - Nothing here exercises the release and snapshot changes. That happens at the next nightly and the next release. CI already builds `-Pspark-4.2` on JDK 17, the JDK both of them use. Separately, `publish_snapshot.yml` has failed every night since at least 2026-09-24 with a 401 from repository.apache.org on its first deploy, so no snapshot is currently being published for any Spark version. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
