[
https://issues.apache.org/jira/browse/HDDS-16174?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Sreeja updated HDDS-16174:
--------------------------
Description:
The users need a way to understand cluster imbalance before starting the
balancer — without changing anything or getting config suggestions. The
assessment command answers: how imbalanced is the cluster, which datanodes are
over/under utilized, and roughly how much data would need to move.
Add ozone admin containerbalancer assessment command which fetches datanode
usage info via the existing getDatanodeUsageInfo RPC, runs the cluster analysis
from HDDS-16173
and shows a human-readable report. It should provide a simple category label
(e.g. large cluster, high imbalance, medium movement ratio) can be derived in
the CLI from the returned summary from analyzer.
Output example:
{panel}
CLUSTER BALANCE ASSESSMENT
═════════════════════════════════════════
Drift: 23.5% (max utilization 91.2% − min utilization 67.7%)
Mean Utilization: 78.4%
Eligible datanodes: 120 datanodes
Category: Large cluster, High imbalance, Medium movement ratio
─── Source Nodes (over-utilized) ─────────────────────
14 datanodes above threshold
Top 5:
dn-worker-017.example.com 91.2% (+13.2% above mean)
dn-worker-004.example.com 89.8% (+11.4% above mean)
dn-worker-022.example.com 88.1% (+9.7% above mean)
dn-worker-011.example.com 87.6% (+9.2% above mean)
dn-worker-039.example.com 86.9% (+8.5% above mean)
─── Target Nodes (under-utilized) ─────────────────────
11 datanodes below threshold
Bottom 5:
dn-worker-118.example.com 67.7% (−10.7% below mean)
dn-worker-119.example.com 68.2% (−10.2% below mean)
dn-worker-120.example.com 69.1% (−9.3% below mean)
dn-worker-115.example.com 70.4% (−8.0% below mean)
dn-worker-116.example.com 71.0% (−7.4% below mean)
─── Movement Summary ─────────────────────────
Total bytes to move: 420 TB
Movement ratio: 3.4% of cluster capacity
{panel}
Threshold defaults to 10% from config, with an optional --threshold override.
Include/exclude datanode lists should also be supported.
Assessment does not show iteration counts, duration estimates, recommended
config, or container move details. Those belong to other commands.
This Jira should also introduce the advisor orchestrator
class(ContainerBalancerAdvisor) with an assess entry point that wires commands
to the shared analyzer class. This class will be later used by other command
that is recommend and dry-run.
was:
The users need a way to understand cluster imbalance before starting the
balancer — without changing anything or getting config suggestions. The
assessment command answers: how imbalanced is the cluster, which datanodes are
over/under utilized, and roughly how much data would need to move.
Add ozone admin containerbalancer assessment command which fetches datanode
usage info via the existing getDatanodeUsageInfo RPC, runs the cluster analysis
from HDDS-16173
and shows a human-readable report. It should provide a simple category label
(e.g. large cluster, high imbalance, medium movement ratio) can be derived in
the CLI from the returned summary from analyzer.
Output example:
{panel}
CLUSTER BALANCE ASSESSMENT
═════════════════════════════════════════
Drift: 23.5% (max utilization 91.2% − min utilization 67.7%)
Mean Utilization: 78.4%
Eligible datanodes: 120 datanodes
Category: Large cluster, High imbalance, Medium movement ratio
─── Source Nodes (over-utilized) ─────────────────────
14 datanodes above threshold
Top 5:
dn-worker-017.example.com 91.2% (+13.2% above mean)
dn-worker-004.example.com 89.8% (+11.4% above mean)
dn-worker-022.example.com 88.1% (+9.7% above mean)
dn-worker-011.example.com 87.6% (+9.2% above mean)
dn-worker-039.example.com 86.9% (+8.5% above mean)
─── Target Nodes (under-utilized) ─────────────────────
11 datanodes below threshold
Bottom 5:
dn-worker-118.example.com 67.7% (−10.7% below mean)
dn-worker-119.example.com 68.2% (−10.2% below mean)
dn-worker-120.example.com 69.1% (−9.3% below mean)
dn-worker-115.example.com 70.4% (−8.0% below mean)
dn-worker-116.example.com 71.0% (−7.4% below mean)
─── Movement Summary ─────────────────────────
Total bytes to move: 420 TB
Movement ratio: 3.4% of cluster capacity
{panel}
Threshold defaults to 10% from config, with an optional --threshold override.
Include/exclude datanode lists should also be supported.
Assessment does not show iteration counts, duration estimates, recommended
config, or container move details. Those belong to other commands.
This Jira should also introduce the advisor orchestrator
class(ContainerBalancerAdvisor) with an assess entry point that wires command
to the shared analyzer class.
> Add container balancer assessment command to report cluster imbalance
> ---------------------------------------------------------------------
>
> Key: HDDS-16174
> URL: https://issues.apache.org/jira/browse/HDDS-16174
> Project: Apache Ozone
> Issue Type: Sub-task
> Reporter: Sreeja
> Assignee: Sravani Revuri
> Priority: Major
>
> The users need a way to understand cluster imbalance before starting the
> balancer — without changing anything or getting config suggestions. The
> assessment command answers: how imbalanced is the cluster, which datanodes
> are over/under utilized, and roughly how much data would need to move.
>
> Add ozone admin containerbalancer assessment command which fetches datanode
> usage info via the existing getDatanodeUsageInfo RPC, runs the cluster
> analysis from HDDS-16173
> and shows a human-readable report. It should provide a simple category label
> (e.g. large cluster, high imbalance, medium movement ratio) can be derived
> in the CLI from the returned summary from analyzer.
> Output example:
> {panel}
> CLUSTER BALANCE ASSESSMENT
> ═════════════════════════════════════════
> Drift: 23.5% (max utilization 91.2% − min utilization 67.7%)
> Mean Utilization: 78.4%
> Eligible datanodes: 120 datanodes
> Category: Large cluster, High imbalance, Medium movement ratio
>
> ─── Source Nodes (over-utilized) ─────────────────────
> 14 datanodes above threshold
> Top 5:
> dn-worker-017.example.com 91.2% (+13.2% above mean)
> dn-worker-004.example.com 89.8% (+11.4% above mean)
> dn-worker-022.example.com 88.1% (+9.7% above mean)
> dn-worker-011.example.com 87.6% (+9.2% above mean)
> dn-worker-039.example.com 86.9% (+8.5% above mean)
>
> ─── Target Nodes (under-utilized) ─────────────────────
> 11 datanodes below threshold
> Bottom 5:
> dn-worker-118.example.com 67.7% (−10.7% below mean)
> dn-worker-119.example.com 68.2% (−10.2% below mean)
> dn-worker-120.example.com 69.1% (−9.3% below mean)
> dn-worker-115.example.com 70.4% (−8.0% below mean)
> dn-worker-116.example.com 71.0% (−7.4% below mean)
>
> ─── Movement Summary ─────────────────────────
> Total bytes to move: 420 TB
> Movement ratio: 3.4% of cluster capacity
>
> {panel}
>
> Threshold defaults to 10% from config, with an optional --threshold override.
> Include/exclude datanode lists should also be supported.
> Assessment does not show iteration counts, duration estimates, recommended
> config, or container move details. Those belong to other commands.
> This Jira should also introduce the advisor orchestrator
> class(ContainerBalancerAdvisor) with an assess entry point that wires
> commands to the shared analyzer class. This class will be later used by other
> command that is recommend and dry-run.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]