[
https://issues.apache.org/jira/browse/HIVE-860?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Ferdinand Xu updated HIVE-860:
------------------------------
Attachment: HIVE-860.2.patch
Reattach the patch to see the CI result. Last time the CI failed to start due
to the following error:
rsync: write failed on
"/data/hive-ptest/logs/PreCommit-HIVE-TRUNK-Build-1868/succeeded/TestCliDriver-avro_add_column.q-orc_wide_table.q-query_with_semi.q-and-12-more/hive.log":
No space left on device (28)
> Persistent distributed cache
> ----------------------------
>
> Key: HIVE-860
> URL: https://issues.apache.org/jira/browse/HIVE-860
> Project: Hive
> Issue Type: Improvement
> Affects Versions: 0.12.0
> Reporter: Zheng Shao
> Assignee: Ferdinand Xu
> Fix For: 0.15.0
>
> Attachments: HIVE-860.1.patch, HIVE-860.2.patch, HIVE-860.2.patch,
> HIVE-860.patch, HIVE-860.patch, HIVE-860.patch, HIVE-860.patch,
> HIVE-860.patch, HIVE-860.patch, HIVE-860.patch, HIVE-860.patch,
> HIVE-860.patch, HIVE-860.patch, HIVE-860.patch
>
>
> DistributedCache is shared across multiple jobs, if the hdfs file name is the
> same.
> We need to make sure Hive put the same file into the same location every time
> and do not overwrite if the file content is the same.
> We can achieve 2 different results:
> A1. Files added with the same name, timestamp, and md5 in the same session
> will have a single copy in distributed cache.
> A2. Filed added with the same name, timestamp, and md5 will have a single
> copy in distributed cache.
> A2 has a bigger benefit in sharing but may raise a question on when Hive
> should clean it up in hdfs.
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)