[ 
https://issues.apache.org/jira/browse/SPARK-59249?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated SPARK-59249:
-----------------------------------
    Labels: pull-request-available  (was: )

> Take the grouped key-row ordering from the shared 
> InternalRowComparableWrapper cache
> ------------------------------------------------------------------------------------
>
>                 Key: SPARK-59249
>                 URL: https://issues.apache.org/jira/browse/SPARK-59249
>             Project: Spark
>          Issue Type: Improvement
>          Components: SQL
>    Affects Versions: 5.0.0
>            Reporter: Peter Toth
>            Priority: Major
>              Labels: pull-request-available
>
> `KeyedPartitioning.groupedKeyRowOrdering` builds the ordering that grouped 
> partition keys are laid out by:
> {code:scala}
> def groupedKeyRowOrdering(dataTypes: Seq[DataType]): BaseOrdering =
>   RowOrdering.createNaturalAscendingOrdering(dataTypes)
> {code}
> That call is byte-for-byte the {{loadFunc}} of 
> {{InternalRowComparableWrapper}}'s private {{orderingCache}}, a 1024-entry 
> {{NonFateSharingCache}}:
> {code:scala}
> private val orderingCache = {
>   val loadFunc = (dataTypes: Seq[DataType]) => {
>     RowOrdering.createNaturalAscendingOrdering(dataTypes)
>   }
>   NonFateSharingCache(loadFunc, MAX_CACHE_ENTRIES)
> }
> {code}
> So every call regenerates an ordering the cache already holds, and that every 
> partition key wrapper already exposes as {{ordering}}. {{GenerateOrdering}} 
> has its own code cache, so nothing is recompiled, but the {{SortOrder}} list 
> and the generated instance are rebuilt each time. The callers are 
> {{KeyedPartitioning.keyRowOrdering}}, 
> {{GroupPartitionsExec.groupAndSortByKeys}} and the reduced-key ordering in 
> {{EnsureRequirements}}.
> The more useful half is not the cost. {{InternalRowComparableWrapper.equals}} 
> is
> {code:scala}
> ordering.compare(row, otherWrapper.row) == 0
> {code}
> over that same ordering, so "two partition keys are equal" and "two partition 
> keys sort together" already come from one definition, written down in two 
> places. Making them share it explicitly is what this ticket is about.
> The proposal is a public 
> {{InternalRowComparableWrapper.orderingFor(dataTypes)}} that returns the 
> cached ordering, with {{groupedKeyRowOrdering}} delegating to it.
> {{DataSourceV2ScanExecBase.outputPartitioning}} is a second site. It calls 
> {{RowOrdering.createNaturalAscendingOrdering}} directly to sort the partition 
> keys it is about to hand to {{KeyedPartitioning}}, which is exactly the 
> grouped key layout order, so it should name that contract rather than rebuild 
> it.
> No behaviour change.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to