[ 
https://issues.apache.org/jira/browse/SPARK-59887?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Peter Toth updated SPARK-59887:
-------------------------------
    Fix Version/s: 4.2.2

> Fix wrong results when a storage-partitioned join pairs a transform of a 
> join-key expression with one of a column
> -----------------------------------------------------------------------------------------------------------------
>
>                 Key: SPARK-59887
>                 URL: https://issues.apache.org/jira/browse/SPARK-59887
>             Project: Spark
>          Issue Type: Bug
>          Components: SQL
>    Affects Versions: 4.2.0, 4.1.3, 4.0.4
>            Reporter: Peter Toth
>            Assignee: Peter Toth
>            Priority: Major
>              Labels: pull-request-available
>             Fix For: 4.3.0, 4.4.0, 4.2.2
>
>
> With {{spark.sql.sources.v2.bucketing.shuffle.enabled}} on, a 
> storage-partitioned join can return wrong results when a side that was 
> shuffled onto a transform of a join-key expression meets another bucketed 
> table in a later join.
> Repro, with AQE and broadcast joins off. {{t1(id bigint)}} and {{t3(x 
> bigint)}} are partitioned by {{bucket(4, ...)}}, {{plain(b bigint)}} is not 
> partitioned, and each holds the values 0 to 7:
> {code:sql}
> SELECT t1.id, p.b, t3.x FROM t1
> JOIN plain p ON t1.id = p.b + 1
> JOIN t3 ON p.b = t3.x
> {code}
> The query should return 7 rows. It returns 0, and the second join has no 
> shuffle.
> Why:
> * The first join shuffles {{plain}} onto {{t1}}'s partitioning. 
> {{KeyedShuffleSpec.createPartitioning}} builds that partitioning from 
> {{plain}}'s join key, so it is {{bucket(4, b + 1)}}.
> * The second join clusters {{plain}} on the bare {{b}}. 
> {{KeyedShuffleSpec.keyPositions}} maps a partition expression to a cluster 
> key by its reference, so {{bucket(4, b + 1)}} counts as a function of {{b}}, 
> the counterpart of {{x}}.
> * {{TransformExpression.isSameFunction}} compares only the function name and 
> the bucket count, so {{bucket(4, b + 1)}} is the same as {{t3}}'s {{bucket(4, 
> x)}}, and the join pairs the partitions as they stand. A row with {{b = x}} 
> sits in bucket {{(b + 1) % 4}} on one side and in {{x % 4}} on the other.
> Measured on master. The one-side shuffle of transform expressions came with 
> SPARK-48012 (4.0.0).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to