[ https://issues.apache.org/jira/browse/HIVE-18263?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16291987#comment-16291987 ]
Hive QA commented on HIVE-18263: -------------------------------- Here are the results of testing the latest attachment: https://issues.apache.org/jira/secure/attachment/12902084/HIVE-18263.1.patch {color:red}ERROR:{color} -1 due to no test(s) being added or modified. {color:red}ERROR:{color} -1 due to 24 failed/errored test(s), 11497 tests executed *Failed tests:* {noformat} TestMiniLlapLocalCliDriver - did not produce a TEST-*.xml file (likely timed out) (batchId=168) [smb_mapjoin_15.q,vector_windowing_expressions.q,auto_join1.q,insert_values_partitioned.q,selectDistinctStar.q,vector_windowing_gby.q,vectorized_timestamp.q,bucket4.q,vector_groupby_mapjoin.q,cbo_subq_exists.q,insert_values_dynamic_partitioned.q,schema_evol_orc_vec_part_all_primitive.q,infer_bucket_sort_bucketed_table.q,cbo_union.q,filter_join_breaktask2.q,udaf_all_keyword.q,schema_evol_orc_vec_part.q,tez_self_join.q,vector_data_types.q,mm_conversions.q,vector_partitioned_date_time.q,schema_evol_orc_nonvec_part.q,dynamic_semijoin_reduction_2.q,vector_mapjoin_reduce.q,vectorization_3.q,auto_sortmerge_join_8.q,disable_merge_for_bucketing.q,vectorized_date_funcs.q,vector_varchar_simple.q,vector_adaptor_usage_mode.q] org.apache.hadoop.hive.cli.TestCliDriver.testCliDriver[ppd_join5] (batchId=35) org.apache.hadoop.hive.cli.TestMiniLlapLocalCliDriver.testCliDriver[bucketsortoptimize_insert_2] (batchId=152) org.apache.hadoop.hive.cli.TestMiniLlapLocalCliDriver.testCliDriver[hybridgrace_hashjoin_2] (batchId=157) org.apache.hadoop.hive.cli.TestMiniLlapLocalCliDriver.testCliDriver[insert_values_orig_table_use_metadata] (batchId=165) org.apache.hadoop.hive.cli.TestMiniLlapLocalCliDriver.testCliDriver[llap_acid] (batchId=169) org.apache.hadoop.hive.cli.TestMiniLlapLocalCliDriver.testCliDriver[llap_acid_fast] (batchId=160) org.apache.hadoop.hive.cli.TestMiniLlapLocalCliDriver.testCliDriver[quotedid_smb] (batchId=157) org.apache.hadoop.hive.cli.TestMiniLlapLocalCliDriver.testCliDriver[sysdb] (batchId=160) org.apache.hadoop.hive.cli.TestNegativeCliDriver.testCliDriver[archive_partspec2] (batchId=93) org.apache.hadoop.hive.cli.TestNegativeCliDriver.testCliDriver[authorization_alter_table_exchange_partition_fail] (batchId=93) org.apache.hadoop.hive.cli.TestNegativeCliDriver.testCliDriver[authorization_part] (batchId=93) org.apache.hadoop.hive.cli.TestNegativeCliDriver.testCliDriver[ctas_noemptyfolder] (batchId=93) org.apache.hadoop.hive.cli.TestNegativeCliDriver.testCliDriver[materialized_view_authorization_create_no_grant] (batchId=93) org.apache.hadoop.hive.cli.TestNegativeCliDriver.testCliDriver[materialized_view_authorization_drop_other] (batchId=93) org.apache.hadoop.hive.cli.TestNegativeCliDriver.testCliDriver[stats_aggregator_error_1] (batchId=93) org.apache.hadoop.hive.cli.TestNegativeCliDriver.testCliDriver[stats_publisher_error_1] (batchId=93) org.apache.hadoop.hive.cli.TestNegativeCliDriver.testCliDriver[subquery_notin_implicit_gby] (batchId=93) org.apache.hadoop.hive.cli.TestNegativeCliDriver.testCliDriver[truncate_bucketed_column] (batchId=93) org.apache.hadoop.hive.cli.TestSparkCliDriver.testCliDriver[auto_sortmerge_join_10] (batchId=138) org.apache.hadoop.hive.cli.TestSparkCliDriver.testCliDriver[bucketsortoptimize_insert_7] (batchId=128) org.apache.hadoop.hive.cli.TestSparkCliDriver.testCliDriver[ppd_join5] (batchId=120) org.apache.hadoop.hive.cli.TestSparkCliDriver.testCliDriver[subquery_multi] (batchId=113) org.apache.hadoop.hive.ql.parse.TestReplicationScenarios.testConstraints (batchId=226) {noformat} Test results: https://builds.apache.org/job/PreCommit-HIVE-Build/8251/testReport Console output: https://builds.apache.org/job/PreCommit-HIVE-Build/8251/console Test logs: http://104.198.109.242/logs/PreCommit-HIVE-Build-8251/ Messages: {noformat} Executing org.apache.hive.ptest.execution.TestCheckPhase Executing org.apache.hive.ptest.execution.PrepPhase Executing org.apache.hive.ptest.execution.YetusPhase Executing org.apache.hive.ptest.execution.ExecutionPhase Executing org.apache.hive.ptest.execution.ReportingPhase Tests exited with: TestsFailedException: 24 tests failed {noformat} This message is automatically generated. ATTACHMENT ID: 12902084 - PreCommit-HIVE-Build > Ptest execution are multiple times slower sometimes due to dying executor > slaves > -------------------------------------------------------------------------------- > > Key: HIVE-18263 > URL: https://issues.apache.org/jira/browse/HIVE-18263 > Project: Hive > Issue Type: Bug > Components: Testing Infrastructure > Reporter: Adam Szita > Assignee: Adam Szita > Attachments: HIVE-18263.0.patch, HIVE-18263.1.patch > > > PreCommit-HIVE-Build job has been seen running very long from time to time. > Usually it should take about 1.5 hours, but in some cases it took over 4-5 > hours. > Looking in the logs of one such execution I've seen that some commands that > were sent to test executing slaves returned 255. Here this typically means > that there is unknown return code for the remote call since hiveptest-server > can't reach these slaves anymore. > In the hiveptest-server logs it is seen that some slaves were killed while > running the job normally, and here is why: > * Hive's ptest-server checks periodically in every 60 minutes the status of > slaves. It also keeps track of slaves that were terminated. > ** If upon such check it is found that a slave that was already killed > ([mTerminatedHosts > map|https://github.com/apache/hive/blob/master/testutils/ptest2/src/main/java/org/apache/hive/ptest/execution/context/CloudExecutionContextProvider.java#L93] > contains its IP) is still running, it will try and terminate it again. > * The server also maintains a file on its local FS that contains the IP of > hosts that were used before. (This probably for resilience reasons) > ** This file is read when tomcat server starts and if any of the IPs in the > file are seen as running slaves, ptest will terminate these first so it can > begin with a fresh start > ** The IPs of these terminated instances already make their way into > {{mTerminatedHosts}} upon initialization... > * The cloud provider may reuse some older IPs, so it is not too rare that the > same IP that belonged to a terminated host is assigned to a new one. > This is problematic: Hive ptest's slave caretaker thread kicks in every 60 > minutes and might see a running host that has the same IP as an old slave had > which was terminated at startup. It will think that this host should be > terminated since it already tried 60 minutes ago as its IP is in > {{mTerminatedHosts}} > We have to fix this by making sure that if a new slave is created, we check > the contents of {{mTerminatedHosts}} and remove this IP from it if it is > there. -- This message was sent by Atlassian JIRA (v6.4.14#64029)