[
https://issues.apache.org/jira/browse/HIVE-2035?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13056205#comment-13056205
]
Siying Dong commented on HIVE-2035:
-----------------------------------
committed
> Use block-level merge for RCFile if merging intermediate results are needed
> ---------------------------------------------------------------------------
>
> Key: HIVE-2035
> URL: https://issues.apache.org/jira/browse/HIVE-2035
> Project: Hive
> Issue Type: Improvement
> Reporter: Ning Zhang
> Assignee: Franklin Hu
> Attachments: hive-2035.1.patch, hive-2035.3.patch
>
>
> Currently if hive.merge.mapredfiles and/or hive.merge.mapfile is set to true
> the intermediate data could be merged using an additional MapReduce job. This
> could be quite expensive if the data size is large. With HIVE-1950, merging
> can be done in the RCFile block level so that it bypasses the
> (de-)compression, (de-)serialization phases. This could improve the merge
> process significantly.
> This JIRA should handle the case where the input table is not stored in
> RCFile, but the destination table is (which requires the intermediate data
> should be stored in the same format as the destination table).
--
This message is automatically generated by JIRA.
For more information on JIRA, see: http://www.atlassian.com/software/jira