Jungtaek Lim created SPARK-60031:
------------------------------------

             Summary: SPIP: External Lookup Join
                 Key: SPARK-60031
                 URL: https://issues.apache.org/jira/browse/SPARK-60031
             Project: Spark
          Issue Type: Epic
          Components: SQL, Structured Streaming
    Affects Versions: 4.4.0
            Reporter: Jungtaek Lim


We want to make stream data to be joined with external storages which are 
(more) efficient to the access pattern of “lookup” based on the key (retrieving 
specific data on demand) instead of loading the data in prior before joining.

For example, suppose we want to join the employee data with department data, to 
show the department name they belong to. (This is called “enrichment”.) Instead 
of loading the whole (or pruned) department data from the external storage at 
once, we read the employee data, and retrieve the department data by associated 
data (e.g. department ID) the employee belongs to, and join the department name 
with employee data.

This will enable us to achieve low latency on join operation since we do not 
wait for loading the data to process but process the stream side immediately 
with querying the static side on demand. This also helps to achieve better 
throughput for the case where Spark cannot derive the efficient filter to prune 
the static side outstandingly, since the new join would only require the data 
“as needed” rather than scanning partial/whole data.

 

SPIP doc: 
[https://docs.google.com/document/d/1fBPM5cku8UJXacG2fw9df1zqALJ-WFm8E9-b9oiFP5M/edit?tab=t.0]



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to