[
https://issues.apache.org/jira/browse/NIFI-16432?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18123873#comment-18123873
]
David Szabo commented on NIFI-16432:
------------------------------------
The [core of
{{ConnectableTask.invoke}}|https://github.com/david-szabo-2/nifi/blob/main/nifi-framework-bundle/nifi-framework/nifi-framework-core/src/main/java/org/apache/nifi/controller/tasks/ConnectableTask.java#L248-L294]
triggering the actual processor code is already surrounded with try-catch
logging any issues (possibly yielding the processor in case of a Throwable)
without rethrowing the exception, which already avoids the problem of the
terminated scheduling chain. The problematic code is in the preceding logic in
the method that decides if the processor should be triggered at all. Most of
that logic seems simple enough (mostly querying the local NiFi) to not warrant
it, except for the call to {{flowController.isPrimary()}} inside
{{[isRunOnCluster|https://github.com/david-szabo-2/nifi/blob/main/nifi-framework-bundle/nifi-framework/nifi-framework-core/src/main/java/org/apache/nifi/controller/tasks/ConnectableTask.java#L125]}}
that may result in (possibly transient) errors depending on the configuration
probably more often during primary election. From the current implementations
of the enclosed {{LeaderElectionManager}} only
{{KubernetesLeaderElectionManager}} may throw exceptions, others are not.
The fix will be two-fold:
* Change {{KubernetesLeaderElectionManager.isLeader()}} so it does not throw
an exception, but returns false instead, aligning with other implementations
* {{flowController.isPrimary()}} to catch exceptions. That is to make sure
even a future {{LeaderElectionManager}} implementation happen to throw an
exception from {{isLeader}} for some reason
> KubernetesLeaderElectionManager errors may stop the scheduling of primary
> only, CRON driven processors
> ------------------------------------------------------------------------------------------------------
>
> Key: NIFI-16432
> URL: https://issues.apache.org/jira/browse/NIFI-16432
> Project: Apache NiFi
> Issue Type: Bug
> Reporter: David Szabo
> Assignee: David Szabo
> Priority: Minor
>
> Cron driven processor executions are scheduled one at a time after a
> successful previous execution. If an exception is thrown at that time the
> processor will not get scheduled again as this scheduling chain is broken.
> Restarting the processors manually resets the scheduling.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)