[
https://issues.apache.org/jira/browse/SPARK-60067?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-60067:
-----------------------------------
Labels: pull-request-available (was: )
> UpdateFields expression size grows with struct width
> ----------------------------------------------------
>
> Key: SPARK-60067
> URL: https://issues.apache.org/jira/browse/SPARK-60067
> Project: Spark
> Issue Type: Bug
> Components: SQL
> Affects Versions: 4.4.0
> Reporter: Ben Hollis
> Priority: Major
> Labels: pull-request-available
>
> Spark represents `Column.withField` and `Column.dropFields` as the Catalyst
> expression `UpdateFields`. Before execution, Spark replaces it with a new
> struct containing one Catalyst expression for every output field, including
> fields that were not changed. A one-field update to an N-field struct
> therefore creates O(N) expression nodes before any row is evaluated. Wide and
> nested structs produce large plans that increase optimizer,
> expression-binding, code-generation, and executor memory costs.
> *Example*
> {code:java}
> val updated = col("wide_struct").withField("field_0", lit(1))
> df.select(updated) {code}
>
> If `wide_struct` has 1,000 fields, this one-field update becomes a
> `CreateNamedStruct` with roughly 1,000 field expressions. The query changes
> one value, but Spark builds and processes a width-sized Catalyst tree before
> execution. Multiple updates or uses repeat that representation and can
> produce plans with tens of thousands of nodes.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]