[
https://issues.apache.org/jira/browse/SPARK-58660?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-58660:
-----------------------------------
Labels: pull-request-available (was: )
> Self-heal the PR Build check when the notify workflow fails to create it
> ------------------------------------------------------------------------
>
> Key: SPARK-58660
> URL: https://issues.apache.org/jira/browse/SPARK-58660
> Project: Spark
> Issue Type: Improvement
> Components: Project Infra
> Affects Versions: 4.3.0
> Reporter: Ganesha S
> Priority: Major
> Labels: pull-request-available
>
> h3. Problem
> A PR's overall CI status on apache/spark is driven by a single check named
> "Build". That check is *created* only by notify_test_workflow.yml, which runs
> once per push via pull_request_target. The scheduled update_build_status.yml
> job (every 15 min) only *updates* an already-existing "Build" check.
> If that one-shot notify job never completes, the "Build" check is never
> created, the scheduled updater has nothing to act on, and the PR is left with
> no Build status (mergeable_state: unstable, combined status: pending) and no
> automatic recovery. The only workaround is for the author to manually re-fire
> a pull_request_target event (push a new commit, or close/reopen the PR).
> The most common trigger is runner starvation: apache/spark shares the ASF's
> limited GitHub Actions capacity, so the notify pull_request_target job can
> sit queued without ever being assigned a runner and then get auto-cancelled.
> In that state the job has an empty runner name, zero executed steps, and an
> empty log archive.
> h3. Example
> [PR #57708|https://github.com/apache/spark/pull/57708], head commit b3d29d1.
> All 33 test check-runs on the fork are green, but the apache-side "Notify
> test workflow" run (31119650844) sat ~45 minutes and was cancelled with no
> steps run, so no "Build" check was ever created. The PR is stuck at
> unstable/pending. Re-running CI in the fork does not help: it only re-runs
> the fork's already-green build and does not re-fire notify_test_workflow.yml,
> and a cancelled apache-side run cannot be re-run by the PR author.
> h3. Root cause
> Check *creation* is a single point of failure (notify, one attempt, no retry,
> no backfill), while only *updates* are on a resilient schedule.
> h3. Fix
> Make the scheduled update_build_status.yml self-healing. While scanning a
> PR's checks it now records whether any "Build" check exists; when none does,
> it backfills one from the forked repo's build_main.yml run for the head SHA,
> mirroring notify_test_workflow.yml. The next scheduled pass then syncs it
> like any other Build check. Recovery becomes automatic within <=15 minutes
> for every PR, with no manual retrigger. notify_test_workflow.yml remains the
> fast path (check appears in seconds on a healthy push); the scheduler is the
> safety net.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]