[
https://issues.apache.org/jira/browse/HBASE-30394?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
huginn updated HBASE-30394:
---------------------------
Component/s: BlockCache
Affects Version/s: 2.4.11
Description:
## What happens
TinyLfuBlockCache reports aggregate cache size and block count through
getCurrentDataSize() and getDataBlockCount(), even though these APIs are
intended to report data-block-only statistics. As a result, data-block metrics
are indistinguishable from aggregate cache metrics when index or metadata
blocks are cached.
## When it happens
This occurs when TinyLfuBlockCache is used and the cache contains a mixture of
data, index, or metadata blocks. The current implementation derives both values
from the aggregate Caffeine cache size and entry count.
## Impact
Operators and monitoring cannot accurately determine the amount and number of
data blocks held by TinyLfuBlockCache, which can make cache composition and
capacity analysis misleading.
## Root cause
On master, TinyLfuBlockCache.getCurrentDataSize() returns getCurrentSize() and
getDataBlockCount() returns getBlockCount(). The insertion and eviction paths
do not maintain separate data-block counters, so the implementation cannot
satisfy the data-block-specific BlockCache contract.
<!-- File:
hbase-server/src/main/java/org/apache/hadoop/hbase/io/hfile/TinyLfuBlockCache.java,
upstream master -->
<!-- Lines: 185-197, 226-232, 296-317, 412-419 -->
## Proposed fix
Maintain the current data-block size and count while blocks are inserted and
removed. Update the counters only for blocks whose BlockType is data, and
return those counters from getCurrentDataSize() and getDataBlockCount().
## Reproduction
Testing evidence will be added by the reporter.
was:
What happens
When ReplicationSourceShipper.clearWALEntryBatch is interrupted while waiting
for the shipper and reader threads to stop, the warning log can leave its final
placeholder unresolved instead of including the interruption value in the
message.
When it happens
During replication source shutdown or termination, if the wait in
clearWALEntryBatch is interrupted before both threads stop.
Impact
The warning is incomplete and makes it harder to identify the interruption that
prevented cleanup from completing.
Root cause
In
hbase-server/src/main/java/org/apache/hadoop/hbase/replication/regionserver/ReplicationSourceShipper.java,
the InterruptedException branch passes the exception as the last argument to a
message with three placeholders. SLF4J treats a final Throwable specially, so
the last placeholder is not populated as intended.
Proposed fix
Format the interrupted exception explicitly with e.toString() for the final
placeholder, preserving the peer ID and thread name in the warning message.
Reproduction
Testing evidence will be added by the reporter.
> Report data block size and count separately in TinyLfuBlockCache
> ----------------------------------------------------------------
>
> Key: HBASE-30394
> URL: https://issues.apache.org/jira/browse/HBASE-30394
> Project: HBase
> Issue Type: Bug
> Components: BlockCache
> Affects Versions: 2.4.11
> Reporter: huginn
> Priority: Major
> Labels: pull-request-available
>
> ## What happens
> TinyLfuBlockCache reports aggregate cache size and block count through
> getCurrentDataSize() and getDataBlockCount(), even though these APIs are
> intended to report data-block-only statistics. As a result, data-block
> metrics are indistinguishable from aggregate cache metrics when index or
> metadata blocks are cached.
> ## When it happens
> This occurs when TinyLfuBlockCache is used and the cache contains a mixture
> of data, index, or metadata blocks. The current implementation derives both
> values from the aggregate Caffeine cache size and entry count.
> ## Impact
> Operators and monitoring cannot accurately determine the amount and number of
> data blocks held by TinyLfuBlockCache, which can make cache composition and
> capacity analysis misleading.
> ## Root cause
> On master, TinyLfuBlockCache.getCurrentDataSize() returns getCurrentSize()
> and getDataBlockCount() returns getBlockCount(). The insertion and eviction
> paths do not maintain separate data-block counters, so the implementation
> cannot satisfy the data-block-specific BlockCache contract.
> <!-- File:
> hbase-server/src/main/java/org/apache/hadoop/hbase/io/hfile/TinyLfuBlockCache.java,
> upstream master -->
> <!-- Lines: 185-197, 226-232, 296-317, 412-419 -->
> ## Proposed fix
> Maintain the current data-block size and count while blocks are inserted and
> removed. Update the counters only for blocks whose BlockType is data, and
> return those counters from getCurrentDataSize() and getDataBlockCount().
> ## Reproduction
> Testing evidence will be added by the reporter.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)