mikamikasuki opened a new pull request, #11366:
URL: https://github.com/apache/arrow-rs/pull/11366

   - Closes #11350
   
   # Rationale for this change
   
   The CSV reader currently recognizes `.` as the decimal mark. Files that use 
a comma decimal mark with a separate field delimiter can be inferred as text or 
fail when read with a floating-point schema.
   
   # What changes are included in this PR?
   
   Add configurable decimal separators to `Format` and `ReaderBuilder`. Schema 
inference, header detection, and floating-point parsing now use the configured 
separator. The default remains `.`, and separators that conflict with CSV 
syntax or floating-point notation are rejected.
   
   # Are these changes tested?
   
   Added regression tests for comma-decimal inference and parsing, the default 
period separator, and invalid separator rejection. The arrow-csv unit and 
doctest suites pass, as do formatting and Clippy checks.
   
   # Are there any significant user-facing changes?
   
   This adds builder methods for configuring the decimal separator; existing 
behavior remains unchanged by default.
   
   # AI assistance
   
   Codex (GPT-6 family) assisted with implementation and regression tests. I 
reviewed the changes and verified them with the arrow-csv test suite, 
formatting, and Clippy.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to