Yaoxuan Wu created SPARK-60078:
----------------------------------
Summary: Column.outer() binds to a column of the inner plan after
a rename, giving wrong results
Key: SPARK-60078
URL: https://issues.apache.org/jira/browse/SPARK-60078
Project: Spark
Issue Type: Bug
Components: SQL
Affects Versions: 4.1.1, 4.2.0, 4.0.0
Reporter: Yaoxuan Wu
If the subquery DataFrame renamed a column ({{{}toDF{}}},
{{{}withColumnRenamed{}}}, {{{}select(col.alias(...)){}}}) and the outer
DataFrame has a column with the old name, an unqualified
{{F.col(old_name).outer()}} is bound to the subquery's own pre-rename column.
The correlation predicate becomes a predicate on the subquery alone and the
query returns wrong results without any error.
{code:java}
from pyspark.sql import SparkSession
spark = SparkSession.builder.getOrCreate()
from pyspark.sql import functions as F
a = spark.createDataFrame([(1,), (2,), (3,)], ['k'])
b = spark.createDataFrame([(1,), (5,)], ['k']).toDF('k_r')
pred = F.col('k_r') == F.col('k').outer() # intended: b.k_r =
a.kprint(a.lateralJoin(b.where(pred)).count())
# 6, expected 1
print(a.where(b.where(pred).exists()).count()) #
3, expected 1
print(a.select(b.where(pred).select(F.count('*')).scalar()).collect()) #
2, 2, 2; expected 1, 0, 0 {code}
Analyzed plan (4.2.0):
{code:java}
LateralJoin lateral-subquery#102 [], Inner
: +- Project [k_r#101L]
: +- Filter (k_r#101L = k#1L)
: +- Project [k#1L AS k_r#101L, k#1L]
: +- LogicalRDD [k#1L], false
+- LogicalRDD [k#0L], false {code}
k is resolved as k#1 of the subquery, not as the outer k#0.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]