[
https://issues.apache.org/jira/browse/CAMEL-25212?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121601#comment-18121601
]
Claus Ibsen commented on CAMEL-25212:
-------------------------------------
Merged to main via https://github.com/apache/camel/pull/27161
> camel-infinispan - InfinispanEmbeddedClusterView and
> InfinispanRemoteClusterView stop refreshing the leadership after one
> exception, and keep the local member as leader when the view is stopped
> -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25212
> URL: https://issues.apache.org/jira/browse/CAMEL-25212
> Project: Camel
> Issue Type: Bug
> Components: camel-infinispan
> Reporter: shashank
> Assignee: shashank
> Priority: Major
> Fix For: 4.23.0
>
>
> Both Infinispan cluster views elect a leader with a {{LeadershipService}}
> task that runs every {{lifespan / 2}} ({{scheduleAtFixedRate}},
> {{InfinispanEmbeddedClusterView}} :166-170, {{InfinispanRemoteClusterView}}
> :174-178): the leader refreshes the leader key with {{replace}} /
> {{replaceWithVersion}}, the others try {{putIfAbsent}}.
> # *One exception ends the election for this node.*
> {{LeadershipService.run()}} (embedded :201-250, remote :213-275) has no
> {{catch}}. Any exception from a cache operation (an Infinispan
> {{TimeoutException}}, an {{AvailabilityException}} in a network partition, a
> HotRod {{TransportException}} while the server cannot be reached, or an
> exception thrown by a leadership listener) propagates out of the task, and
> {{ScheduledThreadPoolExecutor}} then cancels all later runs, silently. From
> then on the node only reacts to removed/expired events of the leader key. The
> leader no longer refreshes the key, but its local member keeps reporting
> {{isLeader() == true}}: when the key expires, another node takes it, and both
> act as leader until the former leader handles the expired event of the key
> (its listener runs the task once, the refresh fails and it steps down). If
> that event does not arrive, the double leadership stays: for example with the
> remote client when the Infinispan server is restarted and the key is lost
> without an event. Afterwards the node only acts on events: a node that is not
> the leader never tries again on its own, and a node that takes the key on an
> event does not refresh it.
> # *Stopping the view keeps the local member as the leader.*
> {{LeadershipService.doStop()}} removes the leader key but never calls
> {{setLeader(false)}}, so no event is fired and
> {{getLocalMember().isLeader()}} stays true. Listeners still registered on the
> view (for example when the view is stopped through JMX, {{stopView}}) keep
> acting as the leader while another node takes the key. This is the problem
> CAMEL-25089 fixed for {{FileLockClusterView}}, which now fires the lost event
> on stop; {{ConsulClusterView}} does too. The ZooKeeper view does not fire it
> on stop since CAMEL-24545 (a deadlock with {{ClusteredRoutePolicy}} at
> shutdown), but it resets its flag; {{ClusteredRoutePolicy}} applies
> leadership changes on its own thread since CAMEL-25062, and
> {{MasterConsumer}} ignores events while it stops.
> h3. Reproduction
> With the embedded view, a {{LOCAL}} cache and a lifespan of 1000 ms (unit
> tests, three runs each):
> * a cache wrapper that throws a {{CacheException}} from the next {{replace}}
> once the node is the leader: afterwards the key is never refreshed again
> ({{replace}} is not called any more within 10 s), while the local member
> still says it is the leader;
> * stopping the view while it is the leader: {{getLocalMember().isLeader()}}
> is still true and no leadership event is fired.
> h3. Proposed fix
> * {{run()}} catches the exception, logs it (WARN, stack trace at DEBUG) and
> gives up the leadership until the next run; the next run takes it again with
> {{putIfAbsent}} if the key is free or still ours ("Lock resumed"). The
> periodic task keeps running.
> * {{doStop()}} gives up the leadership after removing the keys (in a
> {{finally}}), which fires the event.
> * {{LocalMember.setLeader(false)}} fires the event even if looking up the
> current leader fails.
> Tests: {{InfinispanEmbeddedClusterViewLeadershipTest}} (two tests, both fail
> without the fix). The remote view gets the same change; its tests need an
> Infinispan server (Docker), so only its unit tests were run.
> Affected: all versions (no catch in the task at camel-4.0.0, 4.18.0, 4.22.0
> and main).
> Duplicate check (2026-09-30): JIRA text "InfinispanClusterView",
> "InfinispanEmbeddedClusterView", "InfinispanRemoteClusterView", "infinispan
> leadership", component camel-infinispan with "leader"/"cluster"/"master":
> only CAMEL-10287 (original feature), CAMEL-11388 (route policy with a remote
> server), CAMEL-25062 (route policy deadlock). GitHub pull requests
> "infinispan cluster view", "infinispan leadership": none. No open pull
> request touches these files.
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)