[
https://issues.apache.org/jira/browse/SPARK-60045?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-60045:
-----------------------------------
Labels: pull-request-available (was: )
> Avoid per-value regex compilation when parsing decimals in CSV/JSON/XML
> -----------------------------------------------------------------------
>
> Key: SPARK-60045
> URL: https://issues.apache.org/jira/browse/SPARK-60045
> Project: Spark
> Issue Type: Improvement
> Components: SQL
> Affects Versions: 5.0.0
> Reporter: Dustin Smith
> Priority: Major
> Labels: pull-request-available
>
> {{ExprUtils.getDecimalParser}} strips grouping commas from decimal strings
> with {{s.replaceAll(",", "")}} when the locale is {{en-US}} (the default).
> {{String.replaceAll}} compiles a {{java.util.regex.Pattern}} on every call,
> and this parser runs once per decimal value in the CSV reader
> ({{UnivocityParser}}), the JSON reader ({{JacksonParser}}, decimals given as
> strings), the XML reader ({{StaxXmlParser}}) and JSON/XML schema inference.
> The pattern is a literal comma, so {{s.replace(",", "")}} gives exactly the
> same result without a regex.
> Measured on a 2M row CSV file with 5 DECIMAL(18,2) columns (local[1], JDK 17,
> best of 3):
> * one parse: 85.4 ns with {{replaceAll}}, 14.8 ns with {{replace}}
> * reading the file with the DECIMAL schema: 1638 ms before, 842 ms after
> (1.95x)
> * reading the same file with the columns as STRING: 628 ms, so the extra cost
> of decimal parsing drops from about 1010 ms to about 170 ms
> Proposed change: use {{String.replace}} in the {{en-US}} branch of
> {{getDecimalParser}}. No behavior change.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]