[ 
https://issues.apache.org/jira/browse/HDDS-16349?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated HDDS-16349:
----------------------------------
    Labels: pull-request-available  (was: )

> Use HashSet instead of TreeSet for the per-datanode container index
> -------------------------------------------------------------------
>
>                 Key: HDDS-16349
>                 URL: https://issues.apache.org/jira/browse/HDDS-16349
>             Project: Apache Ozone
>          Issue Type: Improvement
>            Reporter: Huang Kuan Hao
>            Assignee: Huang Kuan Hao
>            Priority: Major
>              Labels: pull-request-available
>
> DatanodeEntry keeps each datanode's container set in a TreeSet, so 
> add/remove/contains are O(log N) where N is the number of containers on the 
> datanode (tens of thousands on large clusters). These run per replica on 
> every incremental container report and per container on every full report, so 
> a full report or an SCM-restart registration storm is O(N log N) per datanode.
>   Current:
>   private final Set<ContainerID> containers = new TreeSet<>();
>   public Set<ContainerID> copyContainers() \{ return new 
> TreeSet<>(containers); }
>   DatanodeEntry exposes only Set (no NavigableSet), and no consumer of 
> NodeManager.getContainers(dn) relies on ordering: ContainerReportHandler, 
> DeadNodeHandler, NodeDecommissionManager and DatanodeAdminMonitorImpl 
> iterate/contains, and ContainerBalancerSelectionCriteria copies into its own 
> used-bytes-ordered set. Fix: use HashSmove/contains O(1). Behavior unchanged.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to