Dian-Xuan Yang created HDDS-16359:
-------------------------------------

             Summary: CI dependency resolution fails on Maven Central HTTP 429 
because the configured retries are inert
                 Key: HDDS-16359
                 URL: https://issues.apache.org/jira/browse/HDDS-16359
             Project: Apache Ozone
          Issue Type: Bug
          Components: CI
            Reporter: Dian-Xuan Yang
            Assignee: Dian-Xuan Yang


h2. Problem

CI jobs can fail during dependency resolution, before compilation or tests, 
when Maven Central returns HTTP 429 (Too Many Requests):

{code}
[ERROR] Non-resolvable import POM: ... 
io.opentelemetry:opentelemetry-bom:pom:1.64.0 ...
Could not transfer artifact ... from/to central ... status code: 429
{code}

Example: 
https://github.com/apache/ozone/actions/runs/32979708858/job/99223207224

Two independent factors cause this:

# The Maven cache can miss artifacts when a PR is behind master.

# The configured retry properties are obsolete, so the resolver gives up after 
about 30 seconds.

h2. Why the build contacts Central

All CI workflows restore the Maven repository with:

{code}
key: maven-repo-${{ hashFiles('**/pom.xml') }}
restore-keys: |
maven-repo-
{code}

Only `populate-cache.yml` saves these entries, on pushes to master or release 
branches. A prefix restore therefore gives a cache built from a recent master.

In the failing run, the restored cache was created after master moved 
OpenTelemetry to 1.65.0, while the PR still uses 1.64.0:

{code}
key: maven-repo-cdf51b1427655094fcae42603bbea2137c864cbb8d31bc4e10c35a61c3e9abff
Cache hit for restore-key: 
maven-repo-59b69f0e332d0b30b2a24dc47aa4345f9fd9b73a77d77d6d9e2984fb5ee7fa76
{code}

Therefore, `opentelemetry-bom:1.64.0` is missing and must be downloaded from 
Central.

This is inherent to the current cache scheme: a PR behind master may always 
miss artifacts for versions that master has already replaced. It is not a 
cache-sizing issue.

h2. Why the retry configuration does not help

CI currently sets:

{code}
-Dmaven.wagon.http.retryHandler.class=standard
-Dmaven.wagon.http.retryHandler.count=3
{code}

These were added in 2021, when Maven used Wagon for dependency resolution. The 
runner now uses Maven 3.9.16, which uses `maven-resolver-transport-http` and 
reads `aether.connector.http.retryHandler.*` instead. The old `maven.wagon.*` 
properties are silently ignored.

The resolver defaults are:

{code}
count=3
interval=5000ms
intervalMax=300000ms
serviceUnavailable=429,503
{code}

429 is therefore retried, but with linear backoff:

{code}
0s   request #1 → 429
5s   request #2 → 429
15s  request #3 → 429
30s  request #4 → 429 → fail
{code}

The failing job matches this behavior: the first attempt was around 07:18:53 
and the job failed with 429 at 07:19:25, about 32 seconds later.

The configuration was also reproduced locally with Maven 3.9.16 against a 
repository that always returns 429:

{code}
MAVEN_OPTS                                              requests  elapsed
(none)                                                         4      32s
-Dmaven.wagon.http.retryHandler.count=0                        4      31s  <- 
ignored
-Daether.connector.http.retryHandler.count=0                   1       4s  <- 
honoured
-Daether.connector.http.retryHandler.count=10                 11     268s
{code}

This confirms that the Wagon properties are inert and the `aether` properties 
control the current resolver.

h2. Impact

* CI can fail before running any tests because of a transient Central rate 
limit.
* Contributors cannot re-run failed jobs themselves, requiring a committer 
round trip or force-push.
* These infrastructure failures look like ordinary CI failures and reduce 
confidence in the signal.

h2. Proposed fix

Replace the obsolete Wagon properties with the resolver properties and increase 
the retry count:

{code}
MAVEN_OPTS: -Dhttp.keepAlive=false -Daether.connector.http.retryHandler.count=10
{code}

Apply this to `ci.yml`, `check.yml`, `populate-cache.yml`,
`intermittent-test-check.yml`, `repeat-acceptance.yml`, and `build-ratis.yml`.

This increases retry time from about 30 seconds to about 275 seconds. The 
longest calculated wait is 50 seconds, below the 300-second `intervalMax`, so 
existing `Retry-After` handling remains unchanged. Existing job timeouts also 
prevent an indefinite wait.

`maven.wagon.http.pool` is removed because it is Wagon-specific. 
`http.keepAlive` is a JDK-level property and is therefore left unchanged.

h2. Acceptance criteria

* A PR behind master can tolerate a rate-limited download of a missing artifact 
without failing immediately.
* The CI configuration uses the retry properties actually honored by Maven 
3.9.x.

h2. Out of scope

Maven cache budget cleanup is a separate issue.

Pruning old `maven-repo-*` entries would not help: older PRs may still depend 
on those versions. Restoring only exact keys would also be worse, since a cold 
job would download the entire ~580 MB dependency set instead of a single 
missing artifact.

The current 10 GB cache usage is primarily caused by unrelated entries, 
including eight `Linux-pnpm-*` entries that cannot be restored from their tag 
refs. A stale 495 MB `maven-repo-*` entry also exists on `HDDS-14496-zdu`. 
These should be addressed separately.




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to