[ 
https://issues.apache.org/jira/browse/FLINK-40579?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated FLINK-40579:
-----------------------------------
    Labels: pull-request-available  (was: )

> Weekly connector CI cancels all other branch lines on the first failure and 
> exceeds the ASF 20-concurrent-jobs policy
> ---------------------------------------------------------------------------------------------------------------------
>
>                 Key: FLINK-40579
>                 URL: https://issues.apache.org/jira/browse/FLINK-40579
>             Project: Flink
>          Issue Type: Bug
>          Components: Connectors / Kafka
>    Affects Versions: kafka-5.1.0
>            Reporter: Martijn Visser
>            Assignee: Martijn Visser
>            Priority: Major
>              Labels: pull-request-available
>
> The weekly workflow (.github/workflows/weekly.yml) fans out over five (Flink 
> version, connector
> branch) lines x three JDKs = 30 jobs. Two things make it useless as a signal 
> today:
> 1. Both matrix levels use the GitHub default fail-fast=true. The outer 
> flink_branches matrix in
>    weekly.yml cancels every other branch line ~30 s after the first failing 
> line; the inner JDK
>    matrix in flink-connector-shared-utils ci.yml/python_ci.yml cancels the 
> sibling JDKs ~15 s after
>    the first failing JDK. Every weekly run since at least October 2025 
> therefore shows exactly one
>    Java failure, one Python failure and ~28 cancelled jobs, e.g. run 
> 34002024546 (2026-09-06).
>    Because all 30 jobs start within 2 s and a Java flake surfaces at minute 
> 18-35 of a ~31-minute
>    job, the cancelled jobs have already consumed 90-95 % of their runtime, so 
> fail-fast saves
>    almost no minutes while discarding up to 28 of 30 results. Run 27484280508 
> (2026-06-14) shows the
>    Java matrix fully green while only the Python line made the run red.
> 2. The workflow starts 30 concurrent jobs. 
> https://infra.apache.org/github-actions-policy.html
>    says a workflow MUST have a job concurrency level <= 20 across all 
> matrices and SHOULD be <= 15.
> 3. The pinned Flink patch versions 1.20.4 and 2.1.2 are superseded by 1.20.5 
> and 2.1.3.
> Proposed change (weekly.yml only; the inner per-JDK fail-fast in shared-utils 
> is unchanged):
> - strategy.fail-fast: false on the outer matrix of compile_and_test and 
> python_test.
> - Release lines run one JDK (v5.0/2.2.1 x 17, v5.0/2.1.3 x 17, v4.0/2.0.2 x 
> 17, v3.4/1.20.5 x 11);
>   main/2.2-SNAPSHOT keeps 11/17/21 because push_pr.yml only covers main with 
> 2.2.1. Python runs one
>   JDK per line. 12 concurrent jobs instead of 30, ~280 instead of ~600 runner 
> minutes per run.
>   None of the failures observed in the last 14 weekly runs was JDK dependent.
> - Bump 1.20.4 -> 1.20.5 and 2.1.2 -> 2.1.3.
> - timeout_test: 45 for the Java job (green jobs finish in at most ~36 
> minutes).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to