HippoBaro commented on PR #11262: URL: https://github.com/apache/arrow-rs/pull/11262#issuecomment-5961793835
Thanks @etseidl! Turns out there is a way to unambiguously detect on read wether the input is malformed or not, and I’ve now implemented the reader compatibility you suggested while keeping the writer fix intact: accept both representations on read, but only write the conforming one. The compatibility logic lives in a small private module documenting why these exceptions exist and how they can eventually be removed. Regression coverage uses a single historical fixture, `bad_data/ARROW-RS-GH-11261-FLBA-DICT.parquet`, proposed in apache/parquet-testing#127. It comes from the pre-fix writer and covers V1/V2 pages, nullable and nested values, dictionary fallback, and compression. **This branch currently pins `parquet-testing` to [`1ba58c2971ca`](https://github.com/apache/parquet-testing/commit/1ba58c2971cab87e2337b6185c82a21c219210e3) from my fork, since [parquet-testing#127](https://github.com/apache/parquet-testing/pull/127) has not been merged yet. I’ll update the pin once that's done.** Otherwise I think this one is good to review 🙇 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
