[jira] [Commented] (SPARK-15689) Data source API v2

Reynold Xin (JIRA) Wed, 23 Aug 2017 09:36:26 -0700

    [ 
https://issues.apache.org/jira/browse/SPARK-15689?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16138607#comment-16138607
 ]


Reynold Xin commented on SPARK-15689:
-------------------------------------

Not the author but my guess is that the other approach wouldn't be stable,
neither source nor binary. It relies on internal logical plans.

Note that it would be possible to incorporate the other approach in this as
well, if we add a function that takes in a logical plan and returns a
logical plan.


Wenchen, one thing I have been thinking is whether we should make such user
implementations of data sources immutable. That is, all methods (e.g. Push
filters) would return a new instance of the source. It would be safer. I
don't know if it is worth the complexity though.




> Data source API v2
> ------------------
>
>                 Key: SPARK-15689
>                 URL: https://issues.apache.org/jira/browse/SPARK-15689
>             Project: Spark
>          Issue Type: New Feature
>          Components: SQL
>            Reporter: Reynold Xin
>              Labels: releasenotes
>         Attachments: SPIP Data Source API V2.pdf
>
>
> This ticket tracks progress in creating the v2 of data source API. This new 
> API should focus on:
> 1. Have a small surface so it is easy to freeze and maintain compatibility 
> for a long time. Ideally, this API should survive architectural rewrites and 
> user-facing API revamps of Spark.
> 2. Have a well-defined column batch interface for high performance. 
> Convenience methods should exist to convert row-oriented formats into column 
> batches for data source developers.
> 3. Still support filter push down, similar to the existing API.
> 4. Nice-to-have: support additional common operators, including limit and 
> sampling.
> Note that both 1 and 2 are problems that the current data source API (v1) 
> suffers. The current data source API has a wide surface with dependency on 
> DataFrame/SQLContext, making the data source API compatibility depending on 
> the upper level API. The current data source API is also only row oriented 
> and has to go through an expensive external data type conversion to internal 
> data type.



--
This message was sent by Atlassian JIRA
(v6.4.14#64029)

---------------------------------------------------------------------
To unsubscribe, e-mail: issues-unsubscr...@spark.apache.org
For additional commands, e-mail: issues-h...@spark.apache.org

[jira] [Commented] (SPARK-15689) Data source API v2

Reply via email to