andygrove opened a new issue, #6446:
URL: https://github.com/apache/datafusion-comet/issues/6446
### Describe the bug
Split out of #5149, which fixed the same class of divergence in the string
casts. Related to #3232, whose item 2 covers a different `to_csv` whitespace
problem (trimming before deciding whether to quote).
Comet's `to_csv` implements the `ignoreLeadingWhiteSpace` and
`ignoreTrailingWhiteSpace` write options with Rust's `str::trim_start` /
`str::trim_end` (`native/spark-expr/src/csv_funcs/to_csv.rs`), which strip
Unicode `White_Space`. Spark hands both options to univocity's CSV writer,
whose whitespace is every char `<= ' '`:
- `AbstractWriter.skipLeadingWhitespace` stops at the first char for which
`ch <= ' ' && whitespaceRangeStart < ch` is false, and `WriterCharAppender`
marks trailing chars `<= ' '` as ignored.
- `whitespaceRangeStart` is `-1` because
`CommonSettings.skipBitsAsWhitespace` defaults to `true`, so the set is exactly
`U+0000`-`U+0020`. That is the same set as Java's `String.trim`, and as Comet's
existing `conversion_funcs::trim::trim_java_string` helper.
Both options default to `true` for writing
(`CSVOptions.ignoreLeadingWhiteSpaceFlagInWrite` /
`ignoreTrailingWhiteSpaceFlagInWrite`), so the difference shows with default
options. It goes both ways:
- Control characters `0x00`-`0x08` and `0x0E`-`0x1F` at either end of a
string field are trimmed by Spark and kept by Comet.
- Non-ASCII whitespace such as `U+00A0`, `U+2028` or `U+3000` at either end
of a string field is kept by Spark and trimmed by Comet.
`0x7F` is trimmed by neither.
`to_csv` is `Incompatible` in Comet (#3232), so this only affects queries
that set `spark.comet.expression.StructsToCsv.allowIncompatible=true`.
### Steps to reproduce
With `spark.comet.expression.StructsToCsv.allowIncompatible=true`, over a
Parquet table:
```sql
CREATE TABLE csv_pad(name string, v string) USING parquet;
INSERT INTO csv_pad VALUES
('ctl_0x01', concat(chr(1), 'x', chr(1))),
('ideographic_u3000', concat(cast(X'E38080' as string), 'x',
cast(X'E38080' as string))),
('space', ' x ');
SELECT name, to_csv(named_struct('a', v)) FROM csv_pad ORDER BY name;
```
On the Spark 4.1 profile:
| `name` | Spark | Comet |
| ------------------- | ------- | ----------- |
| `ctl_0x01` | `x` | `\x01x\x01` |
| `ideographic_u3000` | `ăxă` | `x` |
| `space` | `x` | `x` |
### Expected behavior
When `ignoreLeadingWhiteSpace` / `ignoreTrailingWhiteSpace` is set, trim
exactly the chars `<= ' '` from the corresponding end of the value, for example
with leading-only and trailing-only variants of `trim_java_string`.
### Additional context
Found by reading univocity-parsers 2.9.1 (the version Spark bundles) and
Spark 4.1.1's `CSVOptions`, then confirmed by running the query above through
`CometSqlFileTestSuite` with Spark as the oracle.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]