viirya commented on PR #6534:
URL:
https://github.com/apache/datafusion-comet/pull/6534#issuecomment-5966688288
Thanks for the careful read. All three are addressed:
1. `write_sorted_file` now returns `CometError::Internal` when `row_sizes`
is shorter than `row_addresses`, and `columnar_to_row_convert` does the same
when the two address slices differ in length. Both checks run before anything
is read or imported, so a rejected `columnar_to_row_convert` call leaves the
exported structs untouched.
2. The `write_sorted_file` test is now two tests:
- `write_sorted_file_codecs` writes the same rows with `lz4`, `zstd`,
`snappy` and an unknown name, checks the codec tag of every block (the unknown
name gets `LZ4_`), and decodes the blocks back to the input values.
- `write_sorted_file_checksums_the_written_bytes` checks the unseeded
checksum against `crc32fast::hash` of the file, then passes it as
`current_checksum` for a second file and checks the result against a CRC32 of
both files concatenated, which is how it continues across the spill files of a
partition.
3. `shuffle_partition_offsets_after_the_write_is_drained` builds a
`ShuffleWriterExec` over a memory source. It checks the "not drained to
completion" error before execution, then drains the writer and checks that
there are `num_partitions + 1` non-decreasing offsets starting at 0 and ending
at the data file length. `columnar_to_row_convert_imports_and_converts` now has
an `Int32` column and a string column with rows of different lengths, and there
is a test for the mismatched address slices.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]