[jira] [Commented] (BEAM-8306) improve estimation of data byte size reading from source in ElasticsearchIO

Derek He (Jira) Wed, 25 Sep 2019 14:00:40 -0700


    [ 
https://issues.apache.org/jira/browse/BEAM-8306?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16938063#comment-16938063
 ]


Derek He commented on BEAM-8306:
--------------------------------

[~iemejia] pull request [pull 
request|[https://github.com/apache/beam/pull/9660]]

My original feature branch includes fix of BEAM-7916(has been merged into 
master, I guess). But it still show the difference for fix of BEAM-7916.  It 
might make some confusion. To clarify, the fix of  BEAM-8306 is only in 
BoundedElasticsearchSource class in file [ElasticsearchIO.java. 
|https://github.com/apache/beam/pull/9660/files#diff-a826a437a321d9cdf7a37d0bfa35dfe6]

> improve estimation of data byte size reading from source in ElasticsearchIO
> ---------------------------------------------------------------------------
>
>                 Key: BEAM-8306
>                 URL: https://issues.apache.org/jira/browse/BEAM-8306
>             Project: Beam
>          Issue Type: Improvement
>          Components: io-java-elasticsearch
>    Affects Versions: 2.14.0
>            Reporter: Derek He
>            Priority: Major
>
> ElasticsearchIO splits BoundedSource based on the Elasticsearch index size. 
> We expect it can be more accurate to split it base on query result size.
> Currently, we have a big Elasticsearch index. But for query result, it only 
> contains a few documents in the index.  ElasticsearchIO splits it into up 
> to1024 BoundedSources in Google dataflow. It takes long time to finish the 
> processing the small numbers of Elasticsearch document in Google dataflow.
>  
>  



--
This message was sent by Atlassian Jira
(v8.3.4#803005)

[jira] [Commented] (BEAM-8306) improve estimation of data byte size reading from source in ElasticsearchIO

Reply via email to