Hello there,

Following up on our sync status update yesterday, I would like to ask for
some more reviews / verification of my proposed test dataset for the ALP
encoding[1] . It covers a variety of vector sizes and data distributions /
exceptions and I think will help ensure compatibility across
implementations.

Thank you to Curt and Vinoo who have verified their implementations (C# and
Java) can read this file. I have verified it can be read with our Rust
implementation.

I would like to merge this relatively soon as I believe it will unlock full
on implementation in the ecosystem.

Thank you for your time,
Andrew

p.s. Kosta, Prateek and I are also working on a blog post [2] to introduce
this feature and how it works, which hopefully will also accelerate its
rollout


[1]: https://github.com/apache/parquet-testing/pull/119
[2]: https://github.com/apache/parquet-site/pull/195

On Wed, Aug 5, 2026 at 5:49 AM Andrew Lamb <[email protected]> wrote:

> > Just wanted to confirm if you have both low and high precision values
> and
> outliers.
>
> Yes I think the file covers these cases, though it would be great if you
> could double check. The cases are described in more detail in [1]
>
> > Quick question : If we commit the parquet file how does one test that a
> different language reader reads the values correctly?
>
> The first two columns in the file have the same float and double values,
> encoded using PLAIN+zstd. So the verification looks like reading the first
> two columns and comparing them to the in the ALP encoded columns
>
> I think this strategy is more accurate than a CSV as there are no
> potential ambiguity (e.g. exactly what bit pattern does NaN or Inf mean, or
> a decimal value that can't be stored exactly as floating point)
>
> Andrew
>
> [1]:
> https://github.com/alamb/parquet-testing/blob/alamb/alp_test_data/data/README.md#alp-encoding
>
>
> On Tue, Aug 4, 2026 at 9:20 PM PRATEEK GAUR <[email protected]> wrote:
>
>> Thanks Andrew,
>>
>> Yes in our PR [1] we ended up adding all datasets which we needed to prove
>> the value of ALP. (also we didn't trim the number of rows)
>>
>> We tried to cover scenarios like
>> 1) Low precision values
>> 2) High precision values
>> 3) Values with outliers
>> 4) Datasets with different vector sizes
>>
>> And I think you've captured all of that in your dataset.
>> ```
>>
>> message schema {
>>   OPTIONAL FLOAT float_plain;
>>   OPTIONAL DOUBLE double_plain;
>>   OPTIONAL FLOAT float_alp_1024;
>>   OPTIONAL DOUBLE double_alp_1024;
>>   OPTIONAL FLOAT float_alp_4096;
>>   OPTIONAL DOUBLE double_alp_4096;
>>   OPTIONAL FLOAT float_alp_32;
>>   OPTIONAL DOUBLE double_alp_32;
>> }
>> ```
>>
>> Just wanted to confirm if you have both low and high precision values and
>> outliers.
>>
>> Quick question : If we commit the parquet file how does one test that a
>> different language reader reads the values correctly?
>> I thought we would want to write a CSV file too against which we can
>> compare. (or maybe a hash?)
>>
>> [1] : https://github.com/apache/parquet-testing/pull/100
>>
>> On Tue, Aug 4, 2026 at 10:38 AM Andrew Lamb <[email protected]>
>> wrote:
>>
>> > Interop testing is something I am definitely interested in too
>> >
>> > We had a prior discussion[1] and there is an issue about this [2] that
>> > maybe it is time to revive
>> >
>> > [1]: https://lists.apache.org/thread/kd3k4q691lp5c4q3r767zb8jltrm9z33
>> > [2]: https://github.com/apache/parquet-format/issues/441
>> >
>> >
>> > On Tue, Aug 4, 2026 at 1:04 PM Curt Hagenlocher <[email protected]>
>> > wrote:
>> >
>> > > Thanks! I validated that my own implementation
>> > > (https://github.com/clast-project/engineered-wood/pull/64) is working
>> > > with this data.
>> > >
>> > > Better interop testing in general is a recurring topic, and perhaps
>> > > should be taken up again once the versioning work winds down.
>> > >
>> > >
>> > >
>> > > On Tue, Aug 4, 2026 at 8:42 AM Andrew Lamb <[email protected]>
>> > wrote:
>> > > >
>> > > > As part of rolling out ALP to the ecosystem, I think it important to
>> > have
>> > > > an example data set encoded with ALP in parquet-testing for readers
>> to
>> > > test
>> > > > against. Prateek and Vinoo created files like this as part of their
>> > C/C++
>> > > > and Java implementations, but the files were quite large (multiple
>> MB)
>> > > >
>> > > > While working to get the ALP Rust implementation ready to merge[1],
>> I
>> > > spent
>> > > > some time creating a smaller example ALP dataset (211KB) for testing
>> > for
>> > > > consideration[2].
>> > > >
>> > > > I created this file using the ALP C++ implementation and was able to
>> > read
>> > > > it with the ALP Rust implementation.
>> > > >
>> > > > Any feedback would be appreciated,
>> > > > Andrew
>> > > >
>> > > >
>> > > > [1]: https://github.com/apache/arrow-rs/pull/9372
>> > > > [2]: https://github.com/apache/parquet-testing/pull/119
>> > >
>> >
>>
>

Reply via email to