andygrove opened a new pull request, #6417:
URL: https://github.com/apache/datafusion-comet/pull/6417

   ## Which issue does this PR close?
   
   N/A. There is no tracking issue for the promotion.
   
   ## Rationale for this change
   
   Since #4950, Spark 4.2 has the same CI coverage as Spark 3.5 and 4.0: 
Comet's own suites and Apache Spark's SQL suite both run nightly against 
`main`. The docs still call 4.2 experimental and say it should not be used in 
production.
   
   The docs also hid a packaging gap. No Spark 4.2 jar has ever been published: 
neither `dev/release/build-release-comet.sh` nor `publish_snapshot.yml` builds 
the `spark-4.2` profile, and Maven Central has no `comet-spark-spark4.2_2.13` 
artifact. The Gluten comparison's "experimental build also published for the 
Spark 4.2 preview" was never true.
   
   Where 4.2 stands against the promotion criteria in Stage 6 of the 
new-version guide:
   
   - 4.2.0 is a final upstream release.
   - The Spark SQL suite has run nightly only since #4950 merged on 2026-09-28, 
which is short of the guide's one release cycle. The first nightly to include 
it (#6380) failed before running any test, because the 4.2 diff still named 
Comet `1.1.0-SNAPSHOT`. #6398 addresses that, so this PR should merge after 
#6398 lands and a nightly passes.
   - `assume(!isSpark42Plus)` skips remain for #4967 (ANSI overflow error 
tests, labelled `correctness`), #4968 (BloomFilter tests) and #4969 (Iceberg 
REST catalog, which is infrastructure).
   - #4966 (`collect_set` NaN and -0.0 normalization) is open with 
`priority:critical`. #5166 implemented the Spark 4.2 normalization. What 
remains is re-enabling the Spark test in the 4.2 diff, which #5209 tracks.
   
   ## What changes are included in this PR?
   
   User guide and README:
   
   - `installation.md`: 4.2.0 moves into the supported table (Java 17, Scala 
2.13, Nightly / Nightly). The experimental table goes away because 4.2 was its 
only row. The snapshot artifact list and the release download links gain the 
Spark 4.2 jar.
   - `spark-versions.md`: the 4.2 section drops its experimental warning and 
gains a known limitations list: window aggregates with `FILTER (WHERE ...)` 
fall back, a `OneRowRelation` in `UNION` branches takes the union and the 
aggregates above it off Comet (#4949), and the ANSI overflow differences 
(#4967).
   - `README.md` and the Gluten comparison list 4.2 as supported.
   - A Spark 4.2 expression compatibility page: 
`compatibility/expressions/spark-4.2/index.md` plus its toctree entry. #6168 
noted that it was missing.
   
   Packaging and doc generation:
   
   - `dev/release/build-release-comet.sh` and `publish_snapshot.yml` build and 
publish `comet-spark-spark4.2_2.13`, so 1.2.0 ships the jar.
   - `docs/build.sh` and `dev/generate-release-docs.sh` generate the 4.2 pages. 
`spark-4.2` goes before `spark-4.1`, not after it. Every profile run rewrites 
the shared `configs.md` and `expressions.md`, so the last profile wins. Spark 
4.2 registers `max_by` and `min_by` through `MaxByBuilder` and `MinByBuilder`, 
which the generator cannot map to a serde, so with 4.2 last both rows would 
show the unknown placeholder instead of `Native`.
   
   Contributor docs:
   
   - The workflows README and the new-version guide no longer describe 4.2 as 
experimental. Stage 6 of the guide now lists the packaging and doc-generation 
steps above, so the next promotion does not miss them.
   
   Left alone: comments in `pom.xml`, `ci.yml` and `compute-changes.py` still 
call 4.2 experimental. Touching `pom.xml` or `ci.yml` runs the full merge-queue 
suite (macOS, Spark SQL 4.1, Iceberg 1.11 and more) for a comment change, so 
those comments can ride along with the next change to those files. As it 
stands, this PR runs no heavy CI job.
   
   This targets `main` only. `branch-1.1` does not have #4950, including the 
window `FILTER` fallback, so Comet 1.1.0 on Spark 4.2 would evaluate those 
aggregates without their filter. The change should not be backported.
   
   ## How are these changes tested?
   
   - Ran the docs generator for `spark-4.2` the way `docs/build.sh` does. All 
ten category pages generate. Compared with a `spark-4.1` run of the shared 
output: `configs.md` is identical, and `expressions.md` differs only in the 
`max_by` and `min_by` rows described above.
   - Built the docs with Sphinx on this branch and on `main`. The only new 
warnings are for the ten spark-4.2 category pages, which exist only after 
generation. The other versions' pages produce the same warnings in a 
source-only build. With the generated pages copied in, the build has no 
spark-4.2 warnings.
   - `prettier --check` (3.9.9, as in CI), `dev/ci/check-ci-config.py`, `bash 
-n` on the three scripts, and a YAML parse of `publish_snapshot.yml` all pass. 
`compute-changes.py` selects no heavy job for this change set on either 
`pull_request` or `merge_group`.
   - Nothing here exercises the release and snapshot changes. That happens at 
the next nightly and the next release. CI already builds `-Pspark-4.2` on JDK 
17, the JDK both of them use. Separately, `publish_snapshot.yml` has failed 
every night since at least 2026-09-24 with a 401 from repository.apache.org on 
its first deploy, so no snapshot is currently being published for any Spark 
version.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to