akshaytayal opened a new issue, #13139:
URL: https://github.com/apache/gluten/issues/13139
**Umbrella:** GLUTEN-12569 (Spark 4.2 support)
### Background
Spark 4.2 restructured Storage-Partitioned Join (SPJ) partitioning. Verified
against the Spark jars:
- `GroupPartitionsExec` is **new in Spark 4.2.0** (absent in 4.1.1). In 4.2
the scan (`BatchScanExec`) emits **ungrouped** partitions — one per split,
duplicate keys retained — and grouping is deferred to `GroupPartitionsExec`. In
4.1 the scan produced the key-grouped layout itself.
### Problem
The Spark 4.2 shim
(`shims/spark42/src/main/scala/org/apache/spark/sql/execution/datasources/v2/BatchScanExecShim.scala`,
`filteredPartitions`) still rebuilds the old 4.1 layout via
`groupBy(InternalRowComparableWrapper(key))`, collapsing duplicate-key splits.
For keys `[A, A, B]` it advertises 3 partitions but builds only 2 native RDD
partitions; Spark 4.2's `GroupPartitionsExec` then derives parent indices
assuming the ungrouped count → index mismatch / wrong per-key multiplicity.
### Scope / why deferred
Only reachable via scans reporting `KeyGroupedPartitioning` (DSv2
`SupportsReportPartitioning`). Verified `FileScan` (Parquet/ORC/CSV) does
**not** implement it, and Delta does not use DSv2 SPJ. The only native consumer
is the **Iceberg transformer**, which is **not enabled** on Spark 4.2
(dependency unavailable). It is therefore currently **dormant /
non-triggerable** and cannot be behaviorally tested yet.
### Required work (when enabling native keyed DSv2 / Iceberg on Spark 4.2)
1. Preserve one slot per original sorted split (pad filtered-out splits) and
let `GroupPartitionsExec` perform grouping — target layout:
```
original [A1,A2,B1] -> [[A1],[A2],[B1]]
filter removes A2 -> [[A1],[],[B1]]
filter removes A1 and A2 -> [[],[],[B1]]
```
*or* explicitly fall native keyed scans back to vanilla Spark until (1)
is implemented.
2. Add SPJ tests over an **actually offloaded** scan with partition-count
and result assertions (including a nondefault timezone and worker reuse) —
belongs with the `gluten-ut/spark42` UT module.
### Reference
Deferred from review discussion on #13126.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]