Dustin Smith created SPARK-60045:
------------------------------------
Summary: Avoid per-value regex compilation when parsing decimals
in CSV/JSON/XML
Key: SPARK-60045
URL: https://issues.apache.org/jira/browse/SPARK-60045
Project: Spark
Issue Type: Improvement
Components: SQL
Affects Versions: 5.0.0
Reporter: Dustin Smith
{{ExprUtils.getDecimalParser}} strips grouping commas from decimal strings with
{{s.replaceAll(",", "")}} when the locale is {{en-US}} (the default).
{{String.replaceAll}} compiles a {{java.util.regex.Pattern}} on every call, and
this parser runs once per decimal value in the CSV reader
({{UnivocityParser}}), the JSON reader ({{JacksonParser}}, decimals given as
strings), the XML reader ({{StaxXmlParser}}) and JSON/XML schema inference.
The pattern is a literal comma, so {{s.replace(",", "")}} gives exactly the
same result without a regex.
Measured on a 2M row CSV file with 5 DECIMAL(18,2) columns (local[1], JDK 17,
best of 3):
* one parse: 85.4 ns with {{replaceAll}}, 14.8 ns with {{replace}}
* reading the file with the DECIMAL schema: 1638 ms before, 842 ms after (1.95x)
* reading the same file with the columns as STRING: 628 ms, so the extra cost
of decimal parsing drops from about 1010 ms to about 170 ms
Proposed change: use {{String.replace}} in the {{en-US}} branch of
{{getDecimalParser}}. No behavior change.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]