[
https://issues.apache.org/jira/browse/TAJO-1952?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15012842#comment-15012842
]
ASF GitHub Bot commented on TAJO-1952:
--------------------------------------
Github user hyunsik commented on a diff in the pull request:
https://github.com/apache/tajo/pull/846#discussion_r45301565
--- Diff:
tajo-storage/tajo-storage-hdfs/src/main/proto/StorageFragmentProtos.proto ---
@@ -32,3 +32,14 @@ message FileFragmentProto {
repeated string hosts = 5;
repeated int32 disk_ids = 6;
}
+
+message PartitionedFileFragmentProto {
+ required string id = 1;
+ required string path = 2;
+ required int64 start_offset = 3;
+ required int64 length = 4;
+ repeated string hosts = 5;
+ repeated int32 disk_ids = 6;
+ // Partition Name: country=KOREA/city=SEOUL
+ required string partitionName = 7;
--- End diff --
In my opionion, partitionKeys would be more proper because the attribute
includes concatenated partition keys.
> Implement PartitionedFileFragment
> ---------------------------------
>
> Key: TAJO-1952
> URL: https://issues.apache.org/jira/browse/TAJO-1952
> Project: Tajo
> Issue Type: Improvement
> Components: Planner/Optimizer, Storage
> Reporter: Jaehwa Jung
> Assignee: Jaehwa Jung
> Fix For: 0.12.0, 0.11.1
>
> Attachments: TAJO-1952.patch
>
>
> Currently, PartitionedTableScanNode contains the list of partitions and it
> seems to me that the list has some problems as following:
> 1. Duplicate Informs: Task contains Fragment which specify target directory
> or target file for scanning. A path of partition lists already would write to
> Fragment.
> 2. Network Resource: When scanning lost of partition, it will occupy network
> resource, for example, several hundred kilobytes or more. It looks like an
> unnecessary resource because Fragment already has the path of partitions.
> I want to improve above problems by implementing new Fragment called
> PartitionedFileFragment. Currently, I'm planning the implementation as
> following:
> * PartitionedFileFragment will borrow FileFragment and it contains the
> partition path and the partition key values.
> * Remove the path array of partitions from PartitionedTableScanNode.
> * Implement a method for getting filtered partition directories in
> FileTableSpace.
> * Implement a method for making PartitionedFileFragment array.
> * Before making splits, call above method and use it for making splits.
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)