[
https://issues.apache.org/jira/browse/HDDS-16120?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18103890#comment-18103890
]
Chia-Chuan Ho commented on HDDS-16120:
--------------------------------------
Thanks [~smeng] !
The PR is now open: [https://github.com/apache/ozone/pull/10989]
I’d appreciate your review when you have a chance.
> A single MiniOzoneCluster build timeout cascades into whole-suite failures in
> MiniOzoneClusterProvider
> ------------------------------------------------------------------------------------------------------
>
> Key: HDDS-16120
> URL: https://issues.apache.org/jira/browse/HDDS-16120
> Project: Apache Ozone
> Issue Type: Bug
> Reporter: Siyao Meng
> Assignee: Chia-Chuan Ho
> Priority: Major
> Labels: pull-request-available
> Attachments: HDDS-16120.001.patch
>
>
> {{MiniOzoneClusterProvider.createClusters()}} builds clusters on a background
> thread and hands them to consumers through a small blocking queue. When a
> build times out ({{waitForClusterToBeReady()}} throws {{TimeoutException}})
> or fails with {{IOException}}, the background thread rethrows it as an
> unchecked {{RuntimeException("Unable to build cluster")}}. That exception is
> uncaught on the create thread, so the thread dies and the queue is never
> refilled. Every subsequent {{provide()}} then blocks until its own timeout
> and fails with "Failed to obtain available cluster in time".
> The effect is that a single slow or failed cluster build turns into failures
> for every remaining test that shares the provider. This was observed on
> master where one method timed out during cluster startup and the following
> seven methods in the same class all failed with "Failed to obtain available
> cluster in time". It also makes any transient startup slowdown
> disproportionately expensive, since one timeout can push an entire
> integration split to its 90 minute job limit.
> h3. Proposed fix
> Do not let one build failure kill the create thread. Tear down the partial
> cluster and continue (retry) so a single slow build is not fatal to the
> suite, or record the failure and surface the original cause from
> {{provide()}} while keeping the thread alive. Either way removes the "one
> timeout becomes many failures" behavior.
> h3. Testing
> Unit test for {{MiniOzoneClusterProvider}} where the builder throws on one
> build: assert that later {{provide()}} calls still succeed (retry path) or
> fail fast with the original cause, and that the create thread remains alive.
> h3. Note
> This is an amplifier that bounds the blast radius of any single slow cluster
> build; it is valuable independently of whatever triggers a slow build. It is
> not itself the root trigger of the recent master integration-job timeouts,
> and that root trigger is currently unconfirmed.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]