[
https://issues.apache.org/jira/browse/IMPALA-12975?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18106970#comment-18106970
]
ASF subversion and git services commented on IMPALA-12975:
----------------------------------------------------------
Commit b21595276c6c84161eb3194228df016f120ffc17 in impala's branch
refs/heads/master from Joe McDonnell
[ https://gitbox.apache.org/repos/asf?p=impala.git;h=b21595276 ]
IMPALA-15010: Patch gperftools to use size_t for thread cache sizes
This bumps the toolchain, which included a few different fixes.
The primary fix is that gperftools has been patched to use
size_t for thread cache sizes. This would prevent exotic
scenarios where the thread cache size integer could wrap
around to become negative. Google TCMalloc had already made
this change for their thread caching implementation.
The new toolchain also has a different structure for the
layout of the ARM hadoop-client tarball for IMPALA-12975.
When the existing logic in buildall.sh copies the binaries
to the regular HADOOP_HOME, there is a time when the
binaries are partially written that can crash the minicluster.
The original idea for a fix was to leave the ARM binaries
in a different directory and refer to them there.
It turns out to be tedious to get the minicluster to respect
libraries at a separate location from the HADOOP_HOME
location. Instead, this modifies the logic in buildall.sh to
set up a symlink to the ARM hadoop-client location and
only modify it if it is pointing to the wrong place. This
should avoid disrupting a running minicluster.
Testing:
- Ran core jobs on x86_64 and ARM
- Ran a perf-AB-test on ARM
Change-Id: I9dc5e95352f593918f9cacf26786149c8c79b4a2
Reviewed-on: http://gerrit.cloudera.org:8080/24484
Reviewed-by: Joe McDonnell <[email protected]>
Tested-by: Impala Public Jenkins <[email protected]>
Reviewed-by: Michael Smith <[email protected]>
> Rework organization of Hadoop dependency on ARM builds
> ------------------------------------------------------
>
> Key: IMPALA-12975
> URL: https://issues.apache.org/jira/browse/IMPALA-12975
> Project: IMPALA
> Issue Type: Task
> Components: Infrastructure
> Reporter: Joe McDonnell
> Assignee: Joe McDonnell
> Priority: Major
>
> The hadoop binaries that we download from the CDP build number are built for
> x86_64. On x86_64, HADOOP_LIB_DIR and HADOOP_INCLUDE_DIR point to the CDP
> hadoop (i.e. HADOOP_HOME/lib and HADOOP_HOME/include). Various pieces
> (including the C++ build) use these environment variables to find the native
> libraries.
> On ARM, we leave those environment variables pointed to that same location.
> We fix things up by downloading a separate hadoop-client built for ARM, then
> copying the contents into the usual location in the CDP hadoop directory,
> overwriting the x86_64 contents. The code to overwrite the libraries runs on
> each invocation of buildall.sh
> On ARM, we could change this to point HADOOP_LIB_DIR to the downloaded
> hadoop-client (which is built for ARM). With a bit of work on the
> hadoop-client, we could get it to also have the header files and also point
> HADOOP_INCLUDE_DIR to it. This avoids the need to copy files during
> buildall.sh. Any build that wants to pass in a custom hadoop can then use
> HADOOP_LIB_DIR_OVERRIDE and HADOOP_INCLUDE_DIR_OVERRIDE.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]