Djjanks opened a new issue, #11350:
URL: https://github.com/apache/arrow-rs/issues/11350
### Is your feature request related to a problem or challenge?
CSV reader only supports `.` as the decimal separator. Files exported with
European locale settings (e.g. from Excel) use `,` instead, usually with `;` as
the field delimiter, so `1234,5` gets inferred as `Utf8`, and reading with an
explicit `Float64` schema fails with a parse error.
It would be nice to have an option to read and infer floats with `,` as the
decimal separator.
### Describe the solution you'd like
Add a `with_decimal_separator(u8)` option to both `Format` (for schema
inference) and `ReaderBuilder` (for parsing). The default stays `b'.'`, so
existing behaviour doesn't change.
```rust
let format = Format::default()
.with_delimiter(b';')
.with_decimal_separator(b',');
let (schema, _) = format.infer_schema(File::open(PATH).unwrap(),
None).unwrap();
let reader = ReaderBuilder::new(Arc::new(schema))
.with_delimiter(b';')
.with_decimal_separator(b',')
.build(file)
.unwrap();
```
Values that can't work (e.g. the same byte as the delimiter, or a digit)
should be rejected when building the reader.
### Describe alternatives you've considered
Reading these columns as `Utf8` and converting afterwards (e.g.
`replace(',', '.')` + cast).
It works, but schema inference doesn't help here: you have to figure out
yourself which columns are numeric, and it adds an extra pass
### Additional context
Other engines already have this:
- Arrow C++
([`ConvertOptions::decimal_point`](https://arrow.apache.org/docs/cpp/api/formats.html#_CPPv4N5arrow3csv14ConvertOptions13decimal_pointE))
- DuckDB
([`decimal_separator`](https://duckdb.org/docs/current/data/csv/overview#parameters))
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]