Dustin Smith created SPARK-60045:
------------------------------------

             Summary: Avoid per-value regex compilation when parsing decimals 
in CSV/JSON/XML
                 Key: SPARK-60045
                 URL: https://issues.apache.org/jira/browse/SPARK-60045
             Project: Spark
          Issue Type: Improvement
          Components: SQL
    Affects Versions: 5.0.0
            Reporter: Dustin Smith


{{ExprUtils.getDecimalParser}} strips grouping commas from decimal strings with 
{{s.replaceAll(",", "")}} when the locale is {{en-US}} (the default). 
{{String.replaceAll}} compiles a {{java.util.regex.Pattern}} on every call, and 
this parser runs once per decimal value in the CSV reader 
({{UnivocityParser}}), the JSON reader ({{JacksonParser}}, decimals given as 
strings), the XML reader ({{StaxXmlParser}}) and JSON/XML schema inference.

The pattern is a literal comma, so {{s.replace(",", "")}} gives exactly the 
same result without a regex.

Measured on a 2M row CSV file with 5 DECIMAL(18,2) columns (local[1], JDK 17, 
best of 3):
* one parse: 85.4 ns with {{replaceAll}}, 14.8 ns with {{replace}}
* reading the file with the DECIMAL schema: 1638 ms before, 842 ms after (1.95x)
* reading the same file with the columns as STRING: 628 ms, so the extra cost 
of decimal parsing drops from about 1010 ms to about 170 ms

Proposed change: use {{String.replace}} in the {{en-US}} branch of 
{{getDecimalParser}}. No behavior change.




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to