andygrove opened a new issue, #6446:
URL: https://github.com/apache/datafusion-comet/issues/6446

   ### Describe the bug
   
   Split out of #5149, which fixed the same class of divergence in the string 
casts. Related to #3232, whose item 2 covers a different `to_csv` whitespace 
problem (trimming before deciding whether to quote).
   
   Comet's `to_csv` implements the `ignoreLeadingWhiteSpace` and 
`ignoreTrailingWhiteSpace` write options with Rust's `str::trim_start` / 
`str::trim_end` (`native/spark-expr/src/csv_funcs/to_csv.rs`), which strip 
Unicode `White_Space`. Spark hands both options to univocity's CSV writer, 
whose whitespace is every char `<= ' '`:
   
   - `AbstractWriter.skipLeadingWhitespace` stops at the first char for which 
`ch <= ' ' && whitespaceRangeStart < ch` is false, and `WriterCharAppender` 
marks trailing chars `<= ' '` as ignored.
   - `whitespaceRangeStart` is `-1` because 
`CommonSettings.skipBitsAsWhitespace` defaults to `true`, so the set is exactly 
`U+0000`-`U+0020`. That is the same set as Java's `String.trim`, and as Comet's 
existing `conversion_funcs::trim::trim_java_string` helper.
   
   Both options default to `true` for writing 
(`CSVOptions.ignoreLeadingWhiteSpaceFlagInWrite` / 
`ignoreTrailingWhiteSpaceFlagInWrite`), so the difference shows with default 
options. It goes both ways:
   
   - Control characters `0x00`-`0x08` and `0x0E`-`0x1F` at either end of a 
string field are trimmed by Spark and kept by Comet.
   - Non-ASCII whitespace such as `U+00A0`, `U+2028` or `U+3000` at either end 
of a string field is kept by Spark and trimmed by Comet.
   
   `0x7F` is trimmed by neither.
   
   `to_csv` is `Incompatible` in Comet (#3232), so this only affects queries 
that set `spark.comet.expression.StructsToCsv.allowIncompatible=true`.
   
   ### Steps to reproduce
   
   With `spark.comet.expression.StructsToCsv.allowIncompatible=true`, over a 
Parquet table:
   
   ```sql
   CREATE TABLE csv_pad(name string, v string) USING parquet;
   INSERT INTO csv_pad VALUES
     ('ctl_0x01', concat(chr(1), 'x', chr(1))),
     ('ideographic_u3000', concat(cast(X'E38080' as string), 'x', 
cast(X'E38080' as string))),
     ('space', ' x ');
   SELECT name, to_csv(named_struct('a', v)) FROM csv_pad ORDER BY name;
   ```
   
   On the Spark 4.1 profile:
   
   | `name`              | Spark   | Comet       |
   | ------------------- | ------- | ----------- |
   | `ctl_0x01`          | `x`     | `\x01x\x01` |
   | `ideographic_u3000` | ` x ` | `x`         |
   | `space`             | `x`     | `x`         |
   
   ### Expected behavior
   
   When `ignoreLeadingWhiteSpace` / `ignoreTrailingWhiteSpace` is set, trim 
exactly the chars `<= ' '` from the corresponding end of the value, for example 
with leading-only and trailing-only variants of `trim_java_string`.
   
   ### Additional context
   
   Found by reading univocity-parsers 2.9.1 (the version Spark bundles) and 
Spark 4.1.1's `CSVOptions`, then confirmed by running the query above through 
`CometSqlFileTestSuite` with Spark as the oracle.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to