[
https://issues.apache.org/jira/browse/FLINK-38534?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Martijn Visser reopened FLINK-38534:
------------------------------------
Reopening for release-2.2, which has neither df4f8c48d60 nor e009616411c. It
failed in a push run on 2026-09-30 (Java 17, module core):
https://github.com/apache/flink/actions/runs/36725941617/job/109941414612
{code}
[ERROR]
org.apache.flink.runtime.scheduler.adaptive.LocalRecoveryTest.testStateSizeIsConsideredForLocalRecoveryOnRestart
-- Time elapsed: 0.067 s <<< ERROR!
java.util.concurrent.ExecutionException: java.lang.IllegalStateException: Job
0755984428ba5884194741c8e8f506f1 is not a streaming job.
at
java.base/java.util.concurrent.CompletableFuture.reportGet(CompletableFuture.java:396)
at
java.base/java.util.concurrent.CompletableFuture.get(CompletableFuture.java:2073)
at
org.apache.flink.runtime.scheduler.adaptive.AdaptiveSchedulerTestBase.supplyInMainThread(AdaptiveSchedulerTestBase.java:169)
at
org.apache.flink.runtime.scheduler.adaptive.LocalRecoveryTest.testStateSizeIsConsideredForLocalRecoveryOnRestart(LocalRecoveryTest.java:106)
Caused by: java.lang.IllegalStateException: Job
0755984428ba5884194741c8e8f506f1 is not a streaming job.
at
org.apache.flink.runtime.scheduler.adaptive.StateWithExecutionGraph.triggerCheckpoint(StateWithExecutionGraph.java:343)
{code}
This time the fork log has the same deployment collision as the release-2.3
runs:
{code}
14:55:31,712 [ComponentMainThread-pool-1036-thread-1] INFO
org.apache.flink.runtime.executiongraph.ExecutionGraph [] - v1 (1/4)
(a1827b15598ba6cc85f4a94b4d91547e_c578e11f7c4f38761c533198aab3ba5e_0_0)
switched from RUNNING to FAILED on 724555906d9b148c5b673a7a1c5f528f @ localhost
(dataPort=42).
java.lang.IllegalStateException: Cannot deploy v1 (1/4) (attempt #0) with
attempt id
a1827b15598ba6cc85f4a94b4d91547e_c578e11f7c4f38761c533198aab3ba5e_0_0 and
vertex id c578e11f7c4f38761c533198aab3ba5e_0 to
724555906d9b148c5b673a7a1c5f528f @ localhost (dataPort=42) with allocation id
0bf56406659738d22778f5d3c8bbbe56 because execution state has switched to
RUNNING during task restore offload
{code}
Two of the 24 release-2.2 runs in the last seven days hit this test.
release-2.3 has not failed it in the 15 runs since be19c67322a. We should
backport both commits to release-2.2, together with the NPE fix once #29322 is
merged.
> Fix flaky LocalRecoveryTest by waiting for tasks to reach RUNNING state
> -----------------------------------------------------------------------
>
> Key: FLINK-38534
> URL: https://issues.apache.org/jira/browse/FLINK-38534
> Project: Flink
> Issue Type: Bug
> Components: Tests
> Affects Versions: 2.2.0, 2.2.2, 2.3.1
> Reporter: Ruan Hang
> Assignee: mukul mustikar
> Priority: Critical
> Labels: pull-request-available
> Fix For: 2.3.1, 2.4.0
>
>
> {code:java}
> Feb 27 04:21:50 04:21:50.067 [INFO] Results:
> Feb 27 04:21:50 04:21:50.068 [INFO]
> Feb 27 04:21:50 04:21:50.069 [ERROR] Errors:
> Feb 27 04:21:50 04:21:50.070 [ERROR]
> LocalRecoveryTest.testStateSizeIsConsideredForLocalRecoveryOnRestart:113 ยป
> Flink Exhausted retry attempts.
> Feb 27 04:21:50 04:21:50.071 [INFO]
> Feb 27 04:21:50 04:21:50.071 [ERROR] Tests run: 109715, Failures: 0, Errors:
> 1, Skipped: 354
> Feb 27 04:21:50 04:21:50.071 [INFO]
> {code}
> https://dev.azure.com/apache-flink/apache-flink/_build/results?buildId=70334&view=logs&j=77a9d8e1-d610-59b3-fc2a-4766541e0e33&t=25baecb7-cea0-597a-6b01-188b1478210d
--
This message was sent by Atlassian Jira
(v8.20.10#820010)