[ 
https://issues.apache.org/jira/browse/SPARK-58232?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Ganesha S updated SPARK-58232:
------------------------------
    Issue Type: Task  (was: Improvement)

>  [SQL] Simplify cached-batch column-index resolution with a map lookup
> ----------------------------------------------------------------------
>
>                 Key: SPARK-58232
>                 URL: https://issues.apache.org/jira/browse/SPARK-58232
>             Project: Spark
>          Issue Type: Task
>          Components: SQL
>    Affects Versions: 4.2.0
>            Reporter: Ganesha S
>            Priority: Major
>
> *Summary*
> When reading from an in-memory (cached) relation, 
> `DefaultCachedBatchSerializer` maps
> each selected output attribute to its ordinal in the cached schema. It 
> currently does
> this with an `O(n*m)` algorithm that, for every selected attribute, rebuilds 
> the full
> list of cached-schema ExprIds and linear-scans it with indexOf. This is 
> wasteful time
> and transient allocation on wide cached scans and should be a single map 
> lookup.
>  
> *Affected code*
> sql/core/src/main/scala/org/apache/spark/sql/execution/columnar/InMemoryRelation.scala
> Both conversion methods share the pattern:
> * convertCachedBatchToColumnarBatch (vectorized path)
> * convertCachedBatchToInternalRow (row path)
> {code:scala}
> val columnIndices =
>   selectedAttributes.map(a => cacheAttributes.map(o => 
> o.exprId).indexOf(a.exprId)).toArray
> {code}
> For each of m selected attributes this rebuilds a fresh Seq[ExprId] from all 
> n cached
> attributes



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to