[ 
https://issues.apache.org/jira/browse/FLINK-38534?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Martijn Visser reopened FLINK-38534:
------------------------------------

Reopening for release-2.2, which has neither df4f8c48d60 nor e009616411c. It 
failed in a push run on 2026-09-30 (Java 17, module core):
https://github.com/apache/flink/actions/runs/36725941617/job/109941414612

{code}
[ERROR] 
org.apache.flink.runtime.scheduler.adaptive.LocalRecoveryTest.testStateSizeIsConsideredForLocalRecoveryOnRestart
 -- Time elapsed: 0.067 s <<< ERROR!
java.util.concurrent.ExecutionException: java.lang.IllegalStateException: Job 
0755984428ba5884194741c8e8f506f1 is not a streaming job.
        at 
java.base/java.util.concurrent.CompletableFuture.reportGet(CompletableFuture.java:396)
        at 
java.base/java.util.concurrent.CompletableFuture.get(CompletableFuture.java:2073)
        at 
org.apache.flink.runtime.scheduler.adaptive.AdaptiveSchedulerTestBase.supplyInMainThread(AdaptiveSchedulerTestBase.java:169)
        at 
org.apache.flink.runtime.scheduler.adaptive.LocalRecoveryTest.testStateSizeIsConsideredForLocalRecoveryOnRestart(LocalRecoveryTest.java:106)
Caused by: java.lang.IllegalStateException: Job 
0755984428ba5884194741c8e8f506f1 is not a streaming job.
        at 
org.apache.flink.runtime.scheduler.adaptive.StateWithExecutionGraph.triggerCheckpoint(StateWithExecutionGraph.java:343)
{code}

This time the fork log has the same deployment collision as the release-2.3 
runs:

{code}
14:55:31,712 [ComponentMainThread-pool-1036-thread-1] INFO  
org.apache.flink.runtime.executiongraph.ExecutionGraph       [] - v1 (1/4) 
(a1827b15598ba6cc85f4a94b4d91547e_c578e11f7c4f38761c533198aab3ba5e_0_0) 
switched from RUNNING to FAILED on 724555906d9b148c5b673a7a1c5f528f @ localhost 
(dataPort=42).
java.lang.IllegalStateException: Cannot deploy v1 (1/4) (attempt #0) with 
attempt id 
a1827b15598ba6cc85f4a94b4d91547e_c578e11f7c4f38761c533198aab3ba5e_0_0 and 
vertex id c578e11f7c4f38761c533198aab3ba5e_0 to 
724555906d9b148c5b673a7a1c5f528f @ localhost (dataPort=42) with allocation id 
0bf56406659738d22778f5d3c8bbbe56 because execution state has switched to 
RUNNING during task restore offload
{code}

Two of the 24 release-2.2 runs in the last seven days hit this test. 
release-2.3 has not failed it in the 15 runs since be19c67322a. We should 
backport both commits to release-2.2, together with the NPE fix once #29322 is 
merged.

> Fix flaky LocalRecoveryTest by waiting for tasks to reach RUNNING state
> -----------------------------------------------------------------------
>
>                 Key: FLINK-38534
>                 URL: https://issues.apache.org/jira/browse/FLINK-38534
>             Project: Flink
>          Issue Type: Bug
>          Components: Tests
>    Affects Versions: 2.2.0, 2.2.2, 2.3.1
>            Reporter: Ruan Hang
>            Assignee: mukul mustikar
>            Priority: Critical
>              Labels: pull-request-available
>             Fix For: 2.3.1, 2.4.0
>
>
> {code:java}
> Feb 27 04:21:50 04:21:50.067 [INFO] Results:
> Feb 27 04:21:50 04:21:50.068 [INFO] 
> Feb 27 04:21:50 04:21:50.069 [ERROR] Errors: 
> Feb 27 04:21:50 04:21:50.070 [ERROR]   
> LocalRecoveryTest.testStateSizeIsConsideredForLocalRecoveryOnRestart:113 ยป 
> Flink Exhausted retry attempts.
> Feb 27 04:21:50 04:21:50.071 [INFO] 
> Feb 27 04:21:50 04:21:50.071 [ERROR] Tests run: 109715, Failures: 0, Errors: 
> 1, Skipped: 354
> Feb 27 04:21:50 04:21:50.071 [INFO] 
> {code}
> https://dev.azure.com/apache-flink/apache-flink/_build/results?buildId=70334&view=logs&j=77a9d8e1-d610-59b3-fc2a-4766541e0e33&t=25baecb7-cea0-597a-6b01-188b1478210d



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to