Peter Toth created SPARK-59901:
----------------------------------
Summary: Storage-partitioned join fails with INTERNAL_ERROR when a
one-side shuffle's join key references two columns
Key: SPARK-59901
URL: https://issues.apache.org/jira/browse/SPARK-59901
Project: Spark
Issue Type: Bug
Components: SQL
Affects Versions: 5.0.0
Reporter: Peter Toth
With {{spark.sql.sources.v2.bucketing.shuffle.enabled}} on, this fails at
planning with an {{INTERNAL_ERROR}} from the {{assert(refs.size == 1)}} in
{{KeyedShuffleSpec.keyPositions}}. {{b4(id)}} is partitioned by {{bucket(4,
id)}}, and {{pbc(b, c)}} and {{qxy(x, y)}} are not partitioned:
{code:sql}
SELECT * FROM b4
JOIN pbc p ON b4.id = p.b + p.c
JOIN qxy q ON p.b = q.x AND p.c = q.y
{code}
The first join shuffles {{pbc}} onto {{bucket(4, b + c)}}, which
{{KeyedShuffleSpec.createPartitioning}} builds from its join key. The second
join builds a shuffle spec for that layout, and {{keyPositions}} asserts that
each partition expression has exactly one reference. A bucketed third table
fails the same way.
Found while reviewing SPARK-59887. It fails on master and with SPARK-59887
alike.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]