Hello there, Following up on our sync status update yesterday, I would like to ask for some more reviews / verification of my proposed test dataset for the ALP encoding[1] . It covers a variety of vector sizes and data distributions / exceptions and I think will help ensure compatibility across implementations.
Thank you to Curt and Vinoo who have verified their implementations (C# and Java) can read this file. I have verified it can be read with our Rust implementation. I would like to merge this relatively soon as I believe it will unlock full on implementation in the ecosystem. Thank you for your time, Andrew p.s. Kosta, Prateek and I are also working on a blog post [2] to introduce this feature and how it works, which hopefully will also accelerate its rollout [1]: https://github.com/apache/parquet-testing/pull/119 [2]: https://github.com/apache/parquet-site/pull/195 On Wed, Aug 5, 2026 at 5:49 AM Andrew Lamb <[email protected]> wrote: > > Just wanted to confirm if you have both low and high precision values > and > outliers. > > Yes I think the file covers these cases, though it would be great if you > could double check. The cases are described in more detail in [1] > > > Quick question : If we commit the parquet file how does one test that a > different language reader reads the values correctly? > > The first two columns in the file have the same float and double values, > encoded using PLAIN+zstd. So the verification looks like reading the first > two columns and comparing them to the in the ALP encoded columns > > I think this strategy is more accurate than a CSV as there are no > potential ambiguity (e.g. exactly what bit pattern does NaN or Inf mean, or > a decimal value that can't be stored exactly as floating point) > > Andrew > > [1]: > https://github.com/alamb/parquet-testing/blob/alamb/alp_test_data/data/README.md#alp-encoding > > > On Tue, Aug 4, 2026 at 9:20 PM PRATEEK GAUR <[email protected]> wrote: > >> Thanks Andrew, >> >> Yes in our PR [1] we ended up adding all datasets which we needed to prove >> the value of ALP. (also we didn't trim the number of rows) >> >> We tried to cover scenarios like >> 1) Low precision values >> 2) High precision values >> 3) Values with outliers >> 4) Datasets with different vector sizes >> >> And I think you've captured all of that in your dataset. >> ``` >> >> message schema { >> OPTIONAL FLOAT float_plain; >> OPTIONAL DOUBLE double_plain; >> OPTIONAL FLOAT float_alp_1024; >> OPTIONAL DOUBLE double_alp_1024; >> OPTIONAL FLOAT float_alp_4096; >> OPTIONAL DOUBLE double_alp_4096; >> OPTIONAL FLOAT float_alp_32; >> OPTIONAL DOUBLE double_alp_32; >> } >> ``` >> >> Just wanted to confirm if you have both low and high precision values and >> outliers. >> >> Quick question : If we commit the parquet file how does one test that a >> different language reader reads the values correctly? >> I thought we would want to write a CSV file too against which we can >> compare. (or maybe a hash?) >> >> [1] : https://github.com/apache/parquet-testing/pull/100 >> >> On Tue, Aug 4, 2026 at 10:38 AM Andrew Lamb <[email protected]> >> wrote: >> >> > Interop testing is something I am definitely interested in too >> > >> > We had a prior discussion[1] and there is an issue about this [2] that >> > maybe it is time to revive >> > >> > [1]: https://lists.apache.org/thread/kd3k4q691lp5c4q3r767zb8jltrm9z33 >> > [2]: https://github.com/apache/parquet-format/issues/441 >> > >> > >> > On Tue, Aug 4, 2026 at 1:04 PM Curt Hagenlocher <[email protected]> >> > wrote: >> > >> > > Thanks! I validated that my own implementation >> > > (https://github.com/clast-project/engineered-wood/pull/64) is working >> > > with this data. >> > > >> > > Better interop testing in general is a recurring topic, and perhaps >> > > should be taken up again once the versioning work winds down. >> > > >> > > >> > > >> > > On Tue, Aug 4, 2026 at 8:42 AM Andrew Lamb <[email protected]> >> > wrote: >> > > > >> > > > As part of rolling out ALP to the ecosystem, I think it important to >> > have >> > > > an example data set encoded with ALP in parquet-testing for readers >> to >> > > test >> > > > against. Prateek and Vinoo created files like this as part of their >> > C/C++ >> > > > and Java implementations, but the files were quite large (multiple >> MB) >> > > > >> > > > While working to get the ALP Rust implementation ready to merge[1], >> I >> > > spent >> > > > some time creating a smaller example ALP dataset (211KB) for testing >> > for >> > > > consideration[2]. >> > > > >> > > > I created this file using the ALP C++ implementation and was able to >> > read >> > > > it with the ALP Rust implementation. >> > > > >> > > > Any feedback would be appreciated, >> > > > Andrew >> > > > >> > > > >> > > > [1]: https://github.com/apache/arrow-rs/pull/9372 >> > > > [2]: https://github.com/apache/parquet-testing/pull/119 >> > > >> > >> >
