[
https://issues.apache.org/jira/browse/DRILL-6032?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16344294#comment-16344294
]
ASF GitHub Bot commented on DRILL-6032:
---------------------------------------
Github user Ben-Zvi commented on a diff in the pull request:
https://github.com/apache/drill/pull/1101#discussion_r164574746
--- Diff:
exec/java-exec/src/main/java/org/apache/drill/exec/physical/impl/spill/RecordBatchSizer.java
---
@@ -232,9 +251,8 @@ else if (width > 0) {
}
}
- public static final int MAX_VECTOR_SIZE = ValueVector.MAX_BUFFER_SIZE;
// 16 MiB
-
private List<ColumnSize> columnSizes = new ArrayList<>();
+ private Map<String, ColumnSize> columnSizeMap =
CaseInsensitiveMap.newHashMap();
--- End diff --
There may be an issue with MapRDB or HBase which are case sensitive. Also
Drill allows for quoted column names (e.g., "FOO" or "foo") which keep the case
sensitivity.
One crude "fix" can be to add a user option, so when turned on, this code
would use an AbstractHashedMap instead.
> Use RecordBatchSizer to estimate size of columns in HashAgg
> -----------------------------------------------------------
>
> Key: DRILL-6032
> URL: https://issues.apache.org/jira/browse/DRILL-6032
> Project: Apache Drill
> Issue Type: Improvement
> Reporter: Timothy Farkas
> Assignee: Timothy Farkas
> Priority: Major
> Fix For: 1.13.0
>
>
> We need to use the RecordBatchSize to estimate the size of columns in the
> Partition batches created by HashAgg.
--
This message was sent by Atlassian JIRA
(v7.6.3#76005)