[ 
https://issues.apache.org/jira/browse/HDDS-13363?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Ivan Andika updated HDDS-13363:
-------------------------------
    Description: 
Currently, the table cache cleanup is done with using CleanupTableInfo 
annotations on the OMClientResponse. However, this is hard to keep track 
because all the cache updates are done in the OMClientRequest, but we need to 
update the OMClientResponse instead, which is very error-prone. There are a few 
cases where wrong CleanupTableInfo causes memory leak due to uncollected table 
cache entries (I have personally experienced this in our internal cluster). 

OM Table Cache leaks could result in the following issues
 * Gradual increase in OM heap memory: This increase is usually gradual and not 
noticeable since there are no exceptions when a cache entry is not cleaned up.
 * Increased latency in some operations: OM operations like listKeys need to 
iterate both the cache and underlying RocksDB table. Having a lot of cache 
entries might cause increase overall latency.

This silent issue is easily overlooked by contributors and reviewers when 
adding a new OM request/response since it requires deep knowledge of OM double 
buffer mechanism and cache invalidation, which poses high barrier of entry to 
new OM request/response implementation.

This will track improvements to prevent OM table cache leak and also improve 
the current table cache cleanup mechanism.

  was:
Currently, the table cache cleanup is done with using CleanupTableInfo 
annotations on the OMClientResponse. However, this is hard to keep track 
because all the cache updates are done in the OMClientRequest, but we need to 
update the OMClientResponse instead, which is very error-prone. There are a few 
cases where wrong CleanupTableInfo causes memory leak due to uncollected table 
cache entries. 

OM Table Cache leaks could result in the following issues
 * Gradual increase in OM heap memory: This increase is usually gradual and not 
noticeable since there are no exceptions when a cache entry is not cleaned up.
 * Increased latency in some operations: OM operations like listKeys need to 
iterate both the cache and underlying RocksDB table. Having a lot of cache 
entries might cause increase overall latency.

This silent issue is easily overlooked by contributors and reviewers when 
adding a new OM request/response since it requires deep knowledge of OM double 
buffer mechanism and cache invalidation, which poses high barrier of entry to 
new OM request/response implementation.

This will track improvements to prevent OM table cache leak and also improve 
the current table cache cleanup mechanism.


> Prevent OM Table Cache Leak
> ---------------------------
>
>                 Key: HDDS-13363
>                 URL: https://issues.apache.org/jira/browse/HDDS-13363
>             Project: Apache Ozone
>          Issue Type: Improvement
>          Components: OM
>            Reporter: Ivan Andika
>            Assignee: Ivan Andika
>            Priority: Major
>
> Currently, the table cache cleanup is done with using CleanupTableInfo 
> annotations on the OMClientResponse. However, this is hard to keep track 
> because all the cache updates are done in the OMClientRequest, but we need to 
> update the OMClientResponse instead, which is very error-prone. There are a 
> few cases where wrong CleanupTableInfo causes memory leak due to uncollected 
> table cache entries (I have personally experienced this in our internal 
> cluster). 
> OM Table Cache leaks could result in the following issues
>  * Gradual increase in OM heap memory: This increase is usually gradual and 
> not noticeable since there are no exceptions when a cache entry is not 
> cleaned up.
>  * Increased latency in some operations: OM operations like listKeys need to 
> iterate both the cache and underlying RocksDB table. Having a lot of cache 
> entries might cause increase overall latency.
> This silent issue is easily overlooked by contributors and reviewers when 
> adding a new OM request/response since it requires deep knowledge of OM 
> double buffer mechanism and cache invalidation, which poses high barrier of 
> entry to new OM request/response implementation.
> This will track improvements to prevent OM table cache leak and also improve 
> the current table cache cleanup mechanism.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to