andygrove opened a new pull request, #6394:
URL: https://github.com/apache/datafusion-comet/pull/6394

   ## Which issue does this PR close?
   
   Closes #6389.
   
   ## Rationale for this change
   
   A third of the failed merge-queue runs between 2026-09-11 and 2026-09-29 (11 
of 32) failed only on a network fetch, and each one evicted a pull request that 
then had to be re-queued by hand, at the back of the line. Seven of those were 
in the Spark SQL test jobs:
   
   - three where the sbt launcher could not fetch sbt itself, two to four 
seconds into `Run Spark tests` (`[error] [launcher] could not retrieve sbt 
1.11.7`)
   - four where sbt got a `Connection reset` from Maven Central while resolving 
its own dependencies or the build's plugins
   
   The build job's `Pre-compile Spark Test classes` step already retries 
resolution failures, but every test job starts its own sbt and downloads the 
same things again, with no retry. The Iceberg jobs have the same gap on the 
Gradle wrapper's distribution download.
   
   ## What changes are included in this PR?
   
   - `spark_sql_test_reusable.yml`: `Run Spark tests` retries sbt up to twice, 
with the pre-compile step's backoff, but only when the log has a download error 
and no test result line (`[info] - `, `Tests: succeeded`, `Run completed in`). 
A failure after tests start is never retried. The step now runs in bash, like 
the pre-compile step, for `pipefail` and `$RANDOM`.
   - The pre-compile step also retries on `could not retrieve sbt` and `Error 
downloading`.
   - `setup-iceberg-builder` installs the Gradle distribution with `./gradlew 
--version`, retrying up to twice. It runs in all three Iceberg jobs before the 
test step, so a reset or TLS failure fetching Gradle no longer fails a job 
before any test runs.
   
   ## How are these changes tested?
   
   `actionlint` (the one shellcheck note on the step predates this change) and 
`python3 dev/ci/check-ci-config.py` pass locally. I'm applying 
`run-spark-4.1-tests` and `run-iceberg-tests` so both workflows run their 
normal path on this branch before it is queued. The retry path runs only on a 
real download failure, so these runs will not exercise it. The patterns it 
matches come from the failed queue-run logs cited in the issue.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to