Hello,

Three weeks ago we created the following Jira issue FLINK-39989 but there
has been no reaction so far.

 https://issues.apache.org/jira/browse/FLINK-39989

In the issue we described the steps to reproduce a reconciliation failure
in the Flink Kubernetes Operator. This failure arises if the job manager
terminates before the operator has a chance to learn about the job's
termination.

My understanding is that as long as the operator requires observing a job's
termination for correct operation, a finished job must be restored on a
job manager's startup (which is currently not the case).

We would be happy to contribute this change if we can get some guidance
from the dev team.

Thanks,
Filippo

Reply via email to