Peter Toth created SPARK-60043:
----------------------------------
Summary: Keep the sort orders after a partition-key transform in a
V2 scan's reported ordering
Key: SPARK-60043
URL: https://issues.apache.org/jira/browse/SPARK-60043
Project: Spark
Issue Type: Improvement
Components: SQL
Affects Versions: 4.2.0
Reporter: Peter Toth
SPARK-59995 drops every sort order that holds a partition transform from a V2
scan's output ordering. In a reported ordering, such a sort order also ends the
leading run of sort orders that the scan keeps. A sort order on a transform
that is itself a partition key is constant within each partition, like any
partition key, so the sort orders after it still hold. They are dropped anyway:
* keys {{[days(ts)]}}, reported {{[days(ts), id]}}: the scan reports {{[]}}
instead of {{[id]}};
* keys {{[years(ts), id]}}, reported {{[years(ts), id, name]}}: {{[id]}}
instead of {{[id, name]}};
* keys {{[days(ts), id]}}, reported {{[days(ts)]}}: {{[]}}, while no report
would derive {{[id]}}.
Skipping such a sort order instead of ending the run would keep them. Matching
a transform sort order to a transform key needs {{isSameFunction}} plus
semantically equal children rather than {{ExpressionSet}}, since the two are
bound separately and {{BoundFunction.equals}} is only a SHOULD. Falling back to
the derived ordering when nothing of a non-empty report survives would cover
the last case.
This is not a regression, since none of these orderings satisfied a requirement
before SPARK-59995. Measured on the SPARK-59995 PR: two tables that report
{{[years(ts), ts]}} over the key {{years(ts)}}, joined on {{ts}} with
{{preserveOrderingOnCoalesce}} on, get 2 {{SortExec}}s. Skipping the key's sort
order removes both, with correct rows.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]