[
https://issues.apache.org/jira/browse/HBASE-20636?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16491700#comment-16491700
]
Ted Yu commented on HBASE-20636:
--------------------------------
I ran failed tests for patch v2 with patch v1. They also failed.
e.g.
{code}
testMultipleTimestampRanges[3](org.apache.hadoop.hbase.regionserver.TestSeekOptimizations)
Time elapsed: 0.067 s <<< ERROR!
org.apache.hadoop.hbase.DroppedSnapshotException: region:
testMultipleTimestampRanges,,1527348574853.1d8afb02c3e80b1a6ecbd065190b254a.
at
org.apache.hadoop.hbase.regionserver.TestSeekOptimizations.createTimestampRange(TestSeekOptimizations.java:442)
at
org.apache.hadoop.hbase.regionserver.TestSeekOptimizations.testMultipleTimestampRanges(TestSeekOptimizations.java:162)
Caused by: java.lang.IllegalArgumentException: Bloom filter type is ROWPREFIX,
RowPrefixBloomFilter.prefix_length not specified.
at
org.apache.hadoop.hbase.regionserver.TestSeekOptimizations.createTimestampRange(TestSeekOptimizations.java:442)
at
org.apache.hadoop.hbase.regionserver.TestSeekOptimizations.testMultipleTimestampRanges(TestSeekOptimizations.java:162)
{code}
TestSeekOptimizations is parameterized:
{code}
public static final Collection<Object[]> parameters() {
return HBaseTestingUtility.BLOOM_AND_COMPRESSION_COMBINATIONS;
{code}
For the new row prefix filter, filter parameter is needed. Please take this
into account in generating unit test parameter combinations.
> Introduce two bloom filter type : ROWPREFIX and ROWPREFIX_DELIMITED
> -------------------------------------------------------------------
>
> Key: HBASE-20636
> URL: https://issues.apache.org/jira/browse/HBASE-20636
> Project: HBase
> Issue Type: New Feature
> Components: HFile, regionserver, scan
> Reporter: Guangxu Cheng
> Assignee: Guangxu Cheng
> Priority: Major
> Attachments: HBASE-20636.master.001.patch,
> HBASE-20636.master.002.patch
>
>
> As we all know, HBase uses BloomFilter(ROW and ROWCOL) to filter unnecessary
> files to improve read performance. But they only support Get and do not
> support Scan.
> In our company(Tencent), many users need to scan all rows with the same
> prefix, such as Tencent Game. Game user's some operational record will be
> written into HBase, each game user will have a lot of records, the rowkey is
> constructed as userid+'#'+timestamps. So we can scan all records for a given
> user for a specified period.
> For this scenario, we designed the prefix Bloom filter. If the startRow and
> stopRow of the Scan has a valid common prefix, the scan will be allowed to
> use BloomFilter to filter files which will enhance the performance of the
> scan.
> Now, this feature has been running on our cluster over a year, and scan
> performance for this scenario has been improved by more than one times than
> before.
--
This message was sent by Atlassian JIRA
(v7.6.3#76005)