[ 
https://issues.apache.org/jira/browse/CAMEL-25212?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121601#comment-18121601
 ] 

Claus Ibsen commented on CAMEL-25212:
-------------------------------------

Merged to main via https://github.com/apache/camel/pull/27161

> camel-infinispan - InfinispanEmbeddedClusterView and 
> InfinispanRemoteClusterView stop refreshing the leadership after one 
> exception, and keep the local member as leader when the view is stopped
> -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25212
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25212
>             Project: Camel
>          Issue Type: Bug
>          Components: camel-infinispan
>            Reporter: shashank
>            Assignee: shashank
>            Priority: Major
>             Fix For: 4.23.0
>
>
> Both Infinispan cluster views elect a leader with a {{LeadershipService}} 
> task that runs every {{lifespan / 2}} ({{scheduleAtFixedRate}}, 
> {{InfinispanEmbeddedClusterView}} :166-170, {{InfinispanRemoteClusterView}} 
> :174-178): the leader refreshes the leader key with {{replace}} / 
> {{replaceWithVersion}}, the others try {{putIfAbsent}}.
> # *One exception ends the election for this node.* 
> {{LeadershipService.run()}} (embedded :201-250, remote :213-275) has no 
> {{catch}}. Any exception from a cache operation (an Infinispan 
> {{TimeoutException}}, an {{AvailabilityException}} in a network partition, a 
> HotRod {{TransportException}} while the server cannot be reached, or an 
> exception thrown by a leadership listener) propagates out of the task, and 
> {{ScheduledThreadPoolExecutor}} then cancels all later runs, silently. From 
> then on the node only reacts to removed/expired events of the leader key. The 
> leader no longer refreshes the key, but its local member keeps reporting 
> {{isLeader() == true}}: when the key expires, another node takes it, and both 
> act as leader until the former leader handles the expired event of the key 
> (its listener runs the task once, the refresh fails and it steps down). If 
> that event does not arrive, the double leadership stays: for example with the 
> remote client when the Infinispan server is restarted and the key is lost 
> without an event. Afterwards the node only acts on events: a node that is not 
> the leader never tries again on its own, and a node that takes the key on an 
> event does not refresh it.
> # *Stopping the view keeps the local member as the leader.* 
> {{LeadershipService.doStop()}} removes the leader key but never calls 
> {{setLeader(false)}}, so no event is fired and 
> {{getLocalMember().isLeader()}} stays true. Listeners still registered on the 
> view (for example when the view is stopped through JMX, {{stopView}}) keep 
> acting as the leader while another node takes the key. This is the problem 
> CAMEL-25089 fixed for {{FileLockClusterView}}, which now fires the lost event 
> on stop; {{ConsulClusterView}} does too. The ZooKeeper view does not fire it 
> on stop since CAMEL-24545 (a deadlock with {{ClusteredRoutePolicy}} at 
> shutdown), but it resets its flag; {{ClusteredRoutePolicy}} applies 
> leadership changes on its own thread since CAMEL-25062, and 
> {{MasterConsumer}} ignores events while it stops.
> h3. Reproduction
> With the embedded view, a {{LOCAL}} cache and a lifespan of 1000 ms (unit 
> tests, three runs each):
> * a cache wrapper that throws a {{CacheException}} from the next {{replace}} 
> once the node is the leader: afterwards the key is never refreshed again 
> ({{replace}} is not called any more within 10 s), while the local member 
> still says it is the leader;
> * stopping the view while it is the leader: {{getLocalMember().isLeader()}} 
> is still true and no leadership event is fired.
> h3. Proposed fix
> * {{run()}} catches the exception, logs it (WARN, stack trace at DEBUG) and 
> gives up the leadership until the next run; the next run takes it again with 
> {{putIfAbsent}} if the key is free or still ours ("Lock resumed"). The 
> periodic task keeps running.
> * {{doStop()}} gives up the leadership after removing the keys (in a 
> {{finally}}), which fires the event.
> * {{LocalMember.setLeader(false)}} fires the event even if looking up the 
> current leader fails.
> Tests: {{InfinispanEmbeddedClusterViewLeadershipTest}} (two tests, both fail 
> without the fix). The remote view gets the same change; its tests need an 
> Infinispan server (Docker), so only its unit tests were run.
> Affected: all versions (no catch in the task at camel-4.0.0, 4.18.0, 4.22.0 
> and main).
> Duplicate check (2026-09-30): JIRA text "InfinispanClusterView", 
> "InfinispanEmbeddedClusterView", "InfinispanRemoteClusterView", "infinispan 
> leadership", component camel-infinispan with "leader"/"cluster"/"master": 
> only CAMEL-10287 (original feature), CAMEL-11388 (route policy with a remote 
> server), CAMEL-25062 (route policy deadlock). GitHub pull requests 
> "infinispan cluster view", "infinispan leadership": none. No open pull 
> request touches these files.
> _Filed with Claude Code on behalf of allthingssecurity._



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to