[ 
https://issues.apache.org/jira/browse/PHOENIX-7997?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Ishita Dixit reassigned PHOENIX-7997:
-------------------------------------

    Assignee: Ishita Dixit

> Optimize Phoenix MapReduce split generation to lower mapper count
> -----------------------------------------------------------------
>
>                 Key: PHOENIX-7997
>                 URL: https://issues.apache.org/jira/browse/PHOENIX-7997
>             Project: Phoenix
>          Issue Type: Improvement
>            Reporter: Ishita Dixit
>            Assignee: Ishita Dixit
>            Priority: Minor
>
> PhoenixInputFormat currently generates one input split per region, so a 
> MapReduce job over a table launches one mapper per region (#mappers == 
> #regions). For tables with many regions this creates a large number of 
> short-lived mappers, and because multiple regions of a table often reside on 
> the same RegionServer, several mappers hit the same RegionServer 
> concurrently, increasing scheduling overhead and RegionServer hot-spotting.
> Add an opt-in option to PhoenixInputFormat.getSplits() that groups the 
> region-boundary splits by their RegionServer location and emits one combined 
> split per RegionServer (each combined split carrying the underlying 
> per-region scans). This reduces the mapper count from #regions down to 
> approximately #RegionServers and spreads load evenly, one mapper per 
> RegionServer.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to