Jungtaek Lim created SPARK-60031:
------------------------------------
Summary: SPIP: External Lookup Join
Key: SPARK-60031
URL: https://issues.apache.org/jira/browse/SPARK-60031
Project: Spark
Issue Type: Epic
Components: SQL, Structured Streaming
Affects Versions: 4.4.0
Reporter: Jungtaek Lim
We want to make stream data to be joined with external storages which are
(more) efficient to the access pattern of “lookup” based on the key (retrieving
specific data on demand) instead of loading the data in prior before joining.
For example, suppose we want to join the employee data with department data, to
show the department name they belong to. (This is called “enrichment”.) Instead
of loading the whole (or pruned) department data from the external storage at
once, we read the employee data, and retrieve the department data by associated
data (e.g. department ID) the employee belongs to, and join the department name
with employee data.
This will enable us to achieve low latency on join operation since we do not
wait for loading the data to process but process the stream side immediately
with querying the static side on demand. This also helps to achieve better
throughput for the case where Spark cannot derive the efficient filter to prune
the static side outstandingly, since the new join would only require the data
“as needed” rather than scanning partial/whole data.
SPIP doc:
[https://docs.google.com/document/d/1fBPM5cku8UJXacG2fw9df1zqALJ-WFm8E9-b9oiFP5M/edit?tab=t.0]
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]