yuqi1129 opened a new issue, #13459:
URL: https://github.com/apache/gravitino/issues/13459
### What would you like to be improved?
Two CI workflows are close to, or over, their 120-minute job timeout.
Raising `timeout-minutes` again only makes feedback slower. Data from recent
runs (2026-09-20 to 2026-09-22), taken from the `[TEST-TIMING]` log markers:
**`build.yml` / `build (17)`**: successful runs take 87–117 min, and most
recent cancelled PR runs were killed at 120 min.
- `:core:test` alone takes **45 min**. It has Docker-tagged tests, so it
holds `sharedTestEnvironmentLock`, and the other Docker-tagged modules then run
one after another behind it: `hive-metastore-common` 9.8m, `idp-basic` 7.7m,
`iceberg-rest-server`, `paimon`, `doris`, `starrocks`, ...
- `publishToMavenLocal` adds about 8 min in series.
- In the 5 timed-out PR runs I checked, `:core:test` never finished (more
than 105 min), including on PRs that don't touch `core`. There is no per-test
timeout and no thread dump.
**`backend-integration-test.yml`**: `deploy-mysql` and `deploy-postgresql`
take 97–117 min, and `embedded-h2` about 73 min. With `-PskipTests`, every test
task shares the environment lock, so the whole suite runs one test task at a
time: catalog-hive 17.6m, client-java 13.9m, iceberg-rest-server 12.3m,
catalog-fileset 7.7m, catalog-glue 6.8m, paimon 6.8m, lakehouse-iceberg 5.6m,
filesystem-hadoop3 4.4m, hudi 4.1m, others about 12m.
### How should we improve?
Keep one workflow per suite. Inside each workflow, fan out into sharded
sub-jobs through a `shard` matrix dimension, and add one aggregate result job.
**1. Shard definitions in one place**
- Add `dev/ci/test-shards.sh <suite> <shard>`. It prints the Gradle task
arguments for a shard.
- Each suite has one catch-all `others` shard. It is computed by excluding
every explicitly listed shard, so new modules land there automatically and are
never skipped silently.
- Moving a module between shards is a one-line change in this script, with
no workflow YAML edits. The same script can reproduce a shard locally.
**2. `build.yml`**
```
build.yml
├── changes
├── build (matrix: shard = [core, docker-modules, others])
│ core → :core:test (~55m)
│ docker-modules → other Docker-tagged module tests (~35m)
│ others → compile, spotless, checkstyle, non-Docker unit tests,
│ publishToMavenLocal (~30m)
├── coverage (needs: build) merge JaCoCo XML artifacts from all shards
→ PR comment
└── build-result (needs: all, if: always()) single aggregate status check
```
**3. `backend-integration-test.yml`**
```
backend-integration-test.yml
├── changes
├── BackendIT (matrix: backend × test-mode × shard)
│ → backend-integration-test-action.yml (new input: shard)
│ shard = hive catalog-hive, catalog-glue, catalog-lakehouse-hudi
(~29m)
│ client client-java, catalog-fileset, filesystem-hadoop3
(~26m)
│ lakehouse iceberg-rest-server, catalog-lakehouse-iceberg,
│ catalog-lakehouse-paimon, lance-rest-server
(~27m)
│ others everything else
(~12m)
└── BackendIT-result (needs: BackendIT, if: always()) single aggregate
status check
```
- Merge the duplicated `BackendIT-on-push` and `BackendIT-on-pr` jobs into
one matrix job.
- 3 backends × 4 shards = 12 jobs, at about 35–45 min each instead of about
100 min. This stays within the ASF limit of 20 concurrent jobs.
- Name report artifacts by `backend-test-mode-shard` so they don't collide.
**4. Fail fast**
- After the split, lower `timeout-minutes` to about 60 per sub-job.
- Add a default JUnit timeout (`junit.jupiter.execution.timeout.default`)
with a thread dump on timeout.
- Track the `:core:test` hang as a follow-up.
Total runner minutes grow only by the per-shard setup overhead (about 5–8
min per extra job), while PR feedback time drops by about half.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]