[ 
https://issues.apache.org/jira/browse/PHOENIX-7990?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Tanuj Khurana resolved PHOENIX-7990.
------------------------------------
    Resolution: Fixed

> Convergent CAS race in setHAGroupStatusIfNeeded aborts co-active 
> RegionServers on stale-version loss
> ----------------------------------------------------------------------------------------------------
>
>                 Key: PHOENIX-7990
>                 URL: https://issues.apache.org/jira/browse/PHOENIX-7990
>             Project: Phoenix
>          Issue Type: Sub-task
>            Reporter: Tanuj Khurana
>            Assignee: Tanuj Khurana
>            Priority: Major
>
> When two or more active RegionServers react to the same peer-degrade event, 
> each independently attempts to move the shared HA-status znode from 
> ACTIVE_IN_SYNC → ACTIVE_NOT_IN_SYNC via an optimistic compare-and-swap (CAS). 
> The expected version for the CAS comes from the watch-lagged local cache. The 
> first writer wins; the remaining writers CAS against a now-stale version, 
> receive StaleHAGroupStoreRecordVersionException, and abort the RegionServer — 
> even though the shared status has already reached the exact value they 
> intended to write.
> Unnecessary RegionServer aborts during a normal convergent reaction to a peer 
> degrade. The write is idempotent at the group level, so a peer having already 
> set the target state is a success condition, not a conflict — but the code 
> treated it as fatal.
> Root cause
> The validate → build → CAS body ran exactly once against a single 
> cache-derived (record, version) snapshot. Any lost CAS propagated as an 
> exception with no reconciliation against the actual current ZK state.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to