stackedsax commented on issue #47666:
URL: https://github.com/apache/arrow/issues/47666#issuecomment-5981279393
This was fixed by #47741 (GH-47740), released in Arrow 23.0.0, which answers
the open question on #47667 about where the `nullptr` came from.
`DeltaByteArrayDecoderImpl::DecodeArrowDense` zero-initialised a
`std::vector<ByteArray>` (so every entry was `{nullptr, 0}`), `GetInternal()`
legitimately decoded zero values from the corrupted page, and only a
`DCHECK_EQ` compared the expected and actual counts. In release builds that
check is compiled out, so the untouched null entries reached
`ArrowBinaryHelper<FLBAType>::AppendValue`, which memcpy'd `byte_width` bytes
from address 0. #47741 replaced the `DCHECK` with a thrown
`ParquetException("Expected to decode N values, but decoded M values")`, which
rejects the page before any value reaches the helper.
The exact OSS-Fuzz testcase
(`clusterfuzz-testcase-minimized-parquet-arrow-fuzz-4656328221196288`) was
added to arrow-testing in apache/arrow-testing#115 and is present in the
`testing` submodule pinned on main, so it runs as part of the fuzz regression
test in `arrow_reader_writer_test.cc`. On PyArrow 24.0.0 the file now fails
cleanly with `OSError: Expected to decode 2084 values, but decoded 0 values.`
instead of crashing.
Since the root cause is fixed upstream of `AppendValue`, the null guard in
#47667 would no longer be reachable.
@thisisnic I think this can be closed as fixed by #47740, and #47667 closed
along with it.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]