[
https://issues.apache.org/jira/browse/FLINK-40467?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18108705#comment-18108705
]
Josh Keegan commented on FLINK-40467:
-------------------------------------
[~dalelane] no I haven't looked into a fix. We've temporarily worked around it
with a startup script for our app by setting Application ID = Job ID which I
could share, but I expect a proper fix would be in the operator.
> Flink 2.3 Application Mode managed by the kubernetes operator can start
> without submitting the new job after savepoint redeploy
> -------------------------------------------------------------------------------------------------------------------------------
>
> Key: FLINK-40467
> URL: https://issues.apache.org/jira/browse/FLINK-40467
> Project: Flink
> Issue Type: Bug
> Components: Kubernetes Operator
> Affects Versions: 2.3.0
> Reporter: Josh Keegan
> Priority: Minor
>
> With Flink 2.3 Application Mode, we have seen savepoint upgrades complete
> successfully but the replacement application cluster sometimes starts without
> submitting the new job.
> The operator assigns a new Job ID, while the Application ID remains the same
> as the previous deployment. The JobManagers are healthy and REST is
> available, but the operator reports {{{}MISSING / Job Not Found{}}}. Looking
> in the UI, the previous application shows as {{FINISHED}} with no jobs.
> Looking into this, it seems that the operator replaces the Job ID during a
> savepoint redeployment since the old Job has been stopped. The same logic
> doesn't exist for the new Application ID.
> We worked around this by setting the Application ID to the Job ID assigned by
> the operator in a startup script for our docker image. That way, whenever the
> operator rotates the Job ID, our Application ID is also rotated. Since doing
> this approx 1 week ago, we've had no issues.
> It looks like the operator should manage the Application ID alongside the Job
> ID for new Application Mode executions. I think our fix of setting it to be
> the same as the Job ID is fine for single Job deployments, but I don't know
> if more complex logic is required for multiple Jobs.
> Using Flink 2.3 and flink-kubernetes-operator 1.15.0
--
This message was sent by Atlassian Jira
(v8.20.10#820010)