mikamikasuki opened a new pull request, #11366: URL: https://github.com/apache/arrow-rs/pull/11366
- Closes #11350 # Rationale for this change The CSV reader currently recognizes `.` as the decimal mark. Files that use a comma decimal mark with a separate field delimiter can be inferred as text or fail when read with a floating-point schema. # What changes are included in this PR? Add configurable decimal separators to `Format` and `ReaderBuilder`. Schema inference, header detection, and floating-point parsing now use the configured separator. The default remains `.`, and separators that conflict with CSV syntax or floating-point notation are rejected. # Are these changes tested? Added regression tests for comma-decimal inference and parsing, the default period separator, and invalid separator rejection. The arrow-csv unit and doctest suites pass, as do formatting and Clippy checks. # Are there any significant user-facing changes? This adds builder methods for configuring the decimal separator; existing behavior remains unchanged by default. # AI assistance Codex (GPT-6 family) assisted with implementation and regression tests. I reviewed the changes and verified them with the arrow-csv test suite, formatting, and Clippy. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
