[
https://issues.apache.org/jira/browse/HDDS-16132?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Siyao Meng updated HDDS-16132:
------------------------------
Attachment: HDDS-16132.001.patch
> Intermittent timeout in TestSCMFollowerCatchupWithContainerReport (follower
> stuck in safe mode on idle cluster)
> ---------------------------------------------------------------------------------------------------------------
>
> Key: HDDS-16132
> URL: https://issues.apache.org/jira/browse/HDDS-16132
> Project: Apache Ozone
> Issue Type: Sub-task
> Components: SCM HA, test
> Reporter: Siyao Meng
> Priority: Major
> Attachments: HDDS-16132.001.patch
>
>
> {code}
> [ERROR] Tests run: 3, Failures: 0, Errors: 2, Skipped: 0, Time elapsed: 311.2
> s <<< FAILURE! -- in
> org.apache.hadoop.hdds.scm.TestSCMFollowerCatchupWithContainerReport
> [ERROR]
> org.apache.hadoop.hdds.scm.TestSCMFollowerCatchupWithContainerReport.testFollowerCatchupOnIdleCluster
> -- Time elapsed: 125.6 s <<< ERROR!
> java.util.concurrent.TimeoutException:
> Timed out waiting for condition.
> ...
> [ERROR]
> org.apache.hadoop.hdds.scm.TestSCMFollowerCatchupWithContainerReport.testFollowerCatchupAfterContainerCreate
> -- Time elapsed: 121.0 s <<< ERROR!
> java.util.concurrent.TimeoutException:
> {code}
> Both failing methods time out at {{waitFor(() -> !newFollower.isInSafeMode(),
> ...)}} after restarting the follower SCM. In the logs the restarted follower
> stays in safe mode for the full 120s with {{registered datanodes (=0) >=
> required datanodes (=3)}}: no datanode ever registers because the follower's
> {{SCMDatanodeProtocolServer}} is never started.
> Possible cause (needs confirmation from the HDDS-14989 author): HDDS-14989
> defers starting the follower's datanode RPC server to
> {{SCMStateMachine.tryStartDNServerAndRefreshSafeMode()}}, which is invoked
> only from Ratis event callbacks ({{applyTransaction}},
> {{notifyLeaderChanged}}, {{notifyTermIndexUpdated}}) and is gated on
> {{isFollowerCaughtUp()}}. On a restarted follower over an idle cluster,
> {{notifyLeaderChanged}} can fire once while {{getLeaderCommitIndex()}} still
> returns -1, so the DN server is not started; with no further transactions
> there is no later callback to retry, and safe mode never clears. This looks
> like a product-side gap in the idle-follower path rather than pure CI
> flakiness (raising the 120s wait does not help).
> Seen in:
> -
> https://github.com/adoroszlai/ozone-build-results/blob/master/2026/08/08/52075/integration-hdds/hadoop-ozone/integration-test/org.apache.hadoop.hdds.scm.TestSCMFollowerCatchupWithContainerReport.txt
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]