Copilot commented on code in PR #6805:
URL: https://github.com/apache/hive/pull/6805#discussion_r4130999341
##########
ql/src/java/org/apache/hadoop/hive/ql/exec/vector/mapjoin/fast/VectorMapJoinFastHashTableLoader.java:
##########
@@ -312,6 +349,7 @@ public void load(MapJoinTableContainer[] mapJoinTables,
if (!loadExecService.awaitTermination(2, TimeUnit.MINUTES)) {
throw new HiveException("Failed to complete the hash table loader.
Loading timed out.");
}
+ rethrowDrainFailures(loadFutures);
Review Comment:
This check happens only after the producer has read every row and enqueued
the sentinels. If a drain worker fails early, its unbounded
`LinkedBlockingQueue` is never drained, so the producer continues copying and
enqueueing all remaining rows for that partition; a large or skewed input can
exhaust memory or spend the entire load before the original
`MapJoinMemoryExhaustionError` is surfaced. The producer needs to
observe/cancel on the first failed future (or the queues need bounded
failure-aware backpressure).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]