[ 
https://issues.apache.org/jira/browse/SPARK-59901?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Peter Toth updated SPARK-59901:
-------------------------------
    Fix Version/s: 4.2.2

> Storage-partitioned join fails with INTERNAL_ERROR when a one-side shuffle's 
> join key references two columns
> ------------------------------------------------------------------------------------------------------------
>
>                 Key: SPARK-59901
>                 URL: https://issues.apache.org/jira/browse/SPARK-59901
>             Project: Spark
>          Issue Type: Bug
>          Components: SQL
>    Affects Versions: 4.2.0, 4.1.3, 4.0.4
>            Reporter: Peter Toth
>            Assignee: Peter Toth
>            Priority: Major
>             Fix For: 4.3.0, 4.4.0, 4.2.2
>
>
> With {{spark.sql.sources.v2.bucketing.shuffle.enabled}} on, this fails at 
> planning with an {{INTERNAL_ERROR}} from the {{assert(refs.size == 1)}} in 
> {{KeyedShuffleSpec.keyPositions}}. {{b4(id)}} is partitioned by {{bucket(4, 
> id)}}, and {{pbc(b, c)}} and {{qxy(x, y)}} are not partitioned:
> {code:sql}
> SELECT * FROM b4
> JOIN pbc p ON b4.id = p.b + p.c
> JOIN qxy q ON p.b = q.x AND p.c = q.y
> {code}
> The first join shuffles {{pbc}} onto {{bucket(4, b + c)}}, which 
> {{KeyedShuffleSpec.createPartitioning}} builds from its join key. The second 
> join builds a shuffle spec for that layout, and {{keyPositions}} asserts that 
> each partition expression has exactly one reference. A bucketed third table 
> fails the same way.
> Found while reviewing SPARK-59887. It fails on master and with SPARK-59887 
> alike.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to