Peter Toth created SPARK-60044:
----------------------------------
Summary: Freeze the k-way merge ordering of GroupPartitionsExec
when the merge is planned
Key: SPARK-60044
URL: https://issues.apache.org/jira/browse/SPARK-60044
Project: Spark
Issue Type: Bug
Components: SQL
Affects Versions: 4.2.0
Reporter: Peter Toth
{{GroupPartitionsExec}} decides on the k-way merge during planning
({{tryEnableSortedMerge}}), but reads the ordering it merges by
({{kWayMergeOrdering}}) from {{child.outputOrdering}} again when it executes.
The scan's derived part of that ordering follows
{{spark.sql.sources.v2.bucketing.partitionKeyOrdering.enabled}} live. So
turning that conf off after planning leaves a planned merge with an empty
ordering. The comparator then treats every row as equal, and the parent gets
its rows out of order.
Example: a table identity-partitioned by {{(a, b)}} that reports no ordering,
with {{allowKeysSubsetOfPartitionKeys}} and {{preserveOrderingOnCoalesce}} on.
{{SELECT a, b, c, row_number() OVER (PARTITION BY a ORDER BY b) FROM t}} plans
{{GroupPartitionsExec(SortedMerge: true)}} under the window and no
{{SortExec}}. Measured with the in-memory test table: after turning
{{partitionKeyOrdering}} off, the merge's ordering is empty. The rows still
came out in order there, because that source lists its partitions sorted by
key, so the merge happens to keep them in order. A source that lists the splits
of a key in another order would give the parent its rows out of order.
This predates SPARK-59995 and comes with SPARK-56241. It needs a conf change
after planning, like SPARK-59279, which froze the merge config the same way.
The fix would freeze the merge ordering when {{tryEnableSortedMerge}} enables
the merge.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]