[jira] [Commented] (HIVE-17296) Acid tests with multiple splits
[ https://issues.apache.org/jira/browse/HIVE-17296?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16654224#comment-16654224 ] Eugene Koifman commented on HIVE-17296: --- see HIVE-20694 and TestVectorizedOrcAcidRowBatchReader > Acid tests with multiple splits > --- > > Key: HIVE-17296 > URL: https://issues.apache.org/jira/browse/HIVE-17296 > Project: Hive > Issue Type: Test > Components: Transactions >Affects Versions: 3.0.0 >Reporter: Eugene Koifman >Assignee: Eugene Koifman >Priority: Major > > data files in an Acid table are ORC files which may have multiple stripes > for such files in base/ or delta/ (and original files with non acid to acid > conversion) are split by OrcInputFormat into multiple (stripe sized) chunks. > There is additional logic in in OrcRawRecordMerger > (discoverKeyBounds/discoverOriginalKeyBounds) that is not tested by any E2E > tests since none of the have enough data to generate multiple stripes in a > single file. > testRecordReaderOldBaseAndDelta/testRecordReaderNewBaseAndDelta/testOriginalReaderPair > in TestOrcRawRecordMerger has some logic to test this but it really needs e2e > tests. > With ORC-228 it will be possible to write such tests. -- This message was sent by Atlassian JIRA (v7.6.3#76005)
[jira] [Commented] (HIVE-17296) Acid tests with multiple splits
[ https://issues.apache.org/jira/browse/HIVE-17296?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16638973#comment-16638973 ] Eugene Koifman commented on HIVE-17296: --- Hive is now on ORC 1.5.3 > Acid tests with multiple splits > --- > > Key: HIVE-17296 > URL: https://issues.apache.org/jira/browse/HIVE-17296 > Project: Hive > Issue Type: Test > Components: Transactions >Affects Versions: 3.0.0 >Reporter: Eugene Koifman >Assignee: Eugene Koifman >Priority: Blocker > > data files in an Acid table are ORC files which may have multiple stripes > for such files in base/ or delta/ (and original files with non acid to acid > conversion) are split by OrcInputFormat into multiple (stripe sized) chunks. > There is additional logic in in OrcRawRecordMerger > (discoverKeyBounds/discoverOriginalKeyBounds) that is not tested by any E2E > tests since none of the have enough data to generate multiple stripes in a > single file. > testRecordReaderOldBaseAndDelta/testRecordReaderNewBaseAndDelta/testOriginalReaderPair > in TestOrcRawRecordMerger has some logic to test this but it really needs e2e > tests. > With ORC-228 it will be possible to write such tests. -- This message was sent by Atlassian JIRA (v7.6.3#76005)
[jira] [Commented] (HIVE-17296) Acid tests with multiple splits
[ https://issues.apache.org/jira/browse/HIVE-17296?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16127719#comment-16127719 ] Eugene Koifman commented on HIVE-17296: --- ORC-228 is in ORC 1.5. Note that MemoryManager is a ThreadLocal so changing this property may affect other tests. See if this will actually work before backporting > Acid tests with multiple splits > --- > > Key: HIVE-17296 > URL: https://issues.apache.org/jira/browse/HIVE-17296 > Project: Hive > Issue Type: Test > Components: Transactions >Affects Versions: 3.0.0 >Reporter: Eugene Koifman >Assignee: Eugene Koifman >Priority: Critical > > data files in an Acid table are ORC files which may have multiple stripes > for such files in base/ or delta/ (and original files with non acid to acid > conversion) are split by OrcInputFormat into multiple (stripe sized) chunks. > There is additional logic in in OrcRawRecordMerger > (discoverKeyBounds/discoverOriginalKeyBounds) that is not tested by any E2E > tests since none of the have enough data to generate multiple stripes in a > single file. > testRecordReaderOldBaseAndDelta/testRecordReaderNewBaseAndDelta/testOriginalReaderPair > in TestOrcRawRecordMerger has some logic to test this but it really needs e2e > tests. > With ORC-228 it will be possible to write such tests. -- This message was sent by Atlassian JIRA (v6.4.14#64029)