[ 
https://issues.apache.org/jira/browse/HBASE-20636?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16489586#comment-16489586
 ] 

Andrew Purtell commented on HBASE-20636:
----------------------------------------

I think this needs backwards compatibility consideration. It should be possible 
to upgrade to a version with these new BF types and then downgrade. It is fine 
if the downgraded installation will have suboptimal performance until the 
HFiles are rewritten after downgrade via compaction with an older BF type. 

We should never be in a position where we can't read an older HFile version. 
Does this require a HFile version increment? I would argue yes. There should 
also be a fallback. We can at least process the HFile without blooms if we 
don't recognize the bloom filter type.

Likewise, the unknown BF type in the schema should not cause the affected 
regions not to be deployed. 

> Introduce two bloom filter type : ROWPREFIX and ROWPREFIX_DELIMITED
> -------------------------------------------------------------------
>
>                 Key: HBASE-20636
>                 URL: https://issues.apache.org/jira/browse/HBASE-20636
>             Project: HBase
>          Issue Type: New Feature
>          Components: regionserver
>            Reporter: Guangxu Cheng
>            Assignee: Guangxu Cheng
>            Priority: Major
>         Attachments: HBASE-20636.master.001.patch, 
> HBASE-20636.master.002.patch
>
>
> As we all know, HBase uses BloomFilter(ROW and ROWCOL) to filter unnecessary 
> files to improve read performance. But they only support Get and do not 
> support Scan.
> In our company(Tencent), many users need to scan all rows with the same 
> prefix, such as Tencent Game. Game user's some operational record will be 
> written into HBase, each game user will have a lot of records, the rowkey is 
> constructed as userid+'#'+timestamps. So we can scan all records for a given 
> user for a specified period.
> For this scenario, we designed the prefix Bloom filter. If the startRow and 
> stopRow of the Scan has a valid common prefix, the scan will be allowed to 
> use BloomFilter to filter files which will enhance the performance of the 
> scan.
> Now, this feature has been running on our cluster over a year, and scan 
> performance for this scenario has been improved by more than one times than 
> before.



--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

Reply via email to