[
https://issues.apache.org/jira/browse/HBASE-30287?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18108620#comment-18108620
]
mazhengxuan edited comment on HBASE-30287 at 8/27/26 2:59 AM:
--------------------------------------------------------------
Hi [Hernan
Romer|https://issues.apache.org/jira/secure/ViewProfile.jspa?name=hgromer],
I've opened a PR with the fix and regression coverage:
[https://github.com/apache/hbase/pull/8576].
It filters offline and split entries before SnapshotRegionLocator builds its
location maps. The focused regression test covers duplicate start keys, and the
existing split/incremental-backup test passes as well.
When you have some time, could you please help review the changes? Thanks!
was (Author: JIRAUSER298959):
Hi [Hernan
Romer|https://issues.apache.org/jira/secure/ViewProfile.jspa?name=hgromer],
I've opened a PR with the fix and regression coverage:
[https://github.com/apache/hbase/pull/8576].
h5.
h5.
It filters offline and split entries before SnapshotRegionLocator builds its
location maps. The focused regression test covers duplicate start keys, and the
existing split/incremental-backup test passes as well.
When you have some time, could you please help review the changes? Thanks!
> SnapshotRegionLocator should filter out offline regions and split regions
> -------------------------------------------------------------------------
>
> Key: HBASE-30287
> URL: https://issues.apache.org/jira/browse/HBASE-30287
> Project: HBase
> Issue Type: Bug
> Components: backup&restore
> Reporter: Hernan Romer
> Assignee: mazhengxuan
> Priority: Major
> Labels: pull-request-available
>
> I started seeing errors when doing incremental backups
>
> {{2025-09-29 19:31:19.527 [main] INFO org.apache.hadoop.mapreduce.Job -
> Task Id : attempt_1759167518321_0143_m_000000_0, Status : FAILED
> Error: java.lang.IllegalArgumentException: Can't read partitions file
> at
> org.apache.hadoop.mapreduce.lib.partition.TotalOrderPartitioner.setConf(TotalOrderPartitioner.java:117)
> at
> org.apache.hadoop.util.ReflectionUtils.setConf(ReflectionUtils.java:79)
> at
> org.apache.hadoop.util.ReflectionUtils.newInstance(ReflectionUtils.java:140)
> at
> org.apache.hadoop.mapred.MapTask$NewOutputCollector.<init>(MapTask.java:715)
> at org.apache.hadoop.mapred.MapTask.runNewMapper(MapTask.java:783)
> at org.apache.hadoop.mapred.MapTask.run(MapTask.java:348)
> at org.apache.hadoop.mapred.YarnChild$2.run(YarnChild.java:178)
> at
> java.base/java.security.AccessController.doPrivileged(AccessController.java:714)
> at java.base/javax.security.auth.Subject.doAs(Subject.java:525)
> at
> org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1899)
> at org.apache.hadoop.mapred.YarnChild.main(YarnChild.java:172)
> Caused by: java.io.IOException: Wrong number of partitions in keyset
> at
> org.apache.hadoop.mapreduce.lib.partition.TotalOrderPartitioner.setConf(TotalOrderPartitioner.java:91)
> ... 10 more}}
>
> The root cause is that the SnapshotRegionLocator which is used as the
> RegionLocator for modern backups can return dupe start keys in the case of
> region splits.
> {{HFileOutputFormat2}} will do two things
> # Configure the number of reducers based on the # of start keys that we get
> from _all_ region locations
> # De-dupe the start keys and write the partitions based on the de-duped set
> When this happens, the TotalOrderPartitioner fails because it's expecting the
> same number of reducers are partitions.
> We should filter out regions that are either offline, or have been split,
> which is the same thing that the
> [MetaTableAccessor|https://github.com/HubSpot/hbase/blob/70f6120227f9050c8b3cb7c6bb33a768264cf5c4/hbase-client/src/main/java/org/apache/hadoop/hbase/MetaTableAccessor.java#L1261C13-L1261C19]
> does.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)