alamb commented on code in PR #119:
URL: https://github.com/apache/parquet-testing/pull/119#discussion_r3783822869
##########
data/README.md:
##########
@@ -592,6 +593,69 @@ java -jar
parquet-cli/target/parquet-cli-1.16.0-SNAPSHOT-runtime.jar cat /home/r
{"utf8_full_truncation": "Kevin Bacon", "binary_full_truncation": "Kevin
Bacon", "utf8_partial_truncation": "🚀Kevin Bacon", "binary_partial_truncation":
"ÿÿ\u0001\u0002", "utf8_no_truncation": "Ke", "binary_no_truncation": "Ke"}
```
+## ALP encoding
+
+`alp_extended.zstd.parquet` contains FLOAT and DOUBLE columns encoded with
+[Adaptive Lossless floating-Point
(ALP)](https://github.com/apache/parquet-format/blob/master/Encodings.md#adaptive-lossless-floating-point-alp--10)
+(`ALP = 10`).
+It was created with the code in this
[PR](https://github.com/apache/arrow/pull/49154).
+
+All columns contain the same 9032 values. The same values appear all columns
so
+decoders can bit-compare the `ALP` columns against known results stored with
+`PLAIN` encoding.
+
+
+| Column | Encoding | Rationale
/ coverage |
+|-------------------------------------|---------------------------|-----------------------------------------------------------------------------------|
+| `float_plain`, `double_plain` | `PLAIN` + zstd | In-file
reference: readers can bit-compare the ALP columns against these |
+| `float_alp_1024`, `double_alp_1024` | `ALP`, 1024-value vectors | The
default vector size of 1024 values |
+| `float_alp_4096`, `double_alp_4096` | `ALP`, 4096-value vectors | Readers
must honor `log_vector_size` from the page header rather than assume 1024 |
+| `float_alp_32`, `double_alp_32` | `ALP`, 32-value vectors | Many
vectors per page, stresses the per-vector metadata loop |
+
+
+### Data distribution (9032 rows)
+
+The "base distribution" means random values in `[-10.00, 10.00]` with exactly 2
+decimal digits (e.g. `9.43`), which are losslessly encodable by ALP (no
exceptions).
+
+The contents of the 9032 rows are as follows:
+
+| Rows | Contents
| Rationale / coverage
|
+|-----------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
+| 0–1023 | base
| Happy path: full vector, small
frame-of-reference bit width, no exceptions
|
+| 1024–2047 | base, plus: NaN at 1024, 1500 and 2047 (three distinct bit
patterns, see below), +Inf at 2000, −Inf at 2001, −0.0 at 2002, subnormal
(`5e-324` double / `1e-45` float) at 2003 | NaN / Inf and sign/precision edge
values via the exception mechanism; exceptions at exact vector boundaries; NaN
payload preservation; statistics conventions (NaN excluded from min/max, ±Inf
included, `nan_count`, −0.0/+0.0 normalization) |
Review Comment:
I did not mean to convolve the idea of nan statistics with the rest of this
file (even though it contains float/double data, I was trying to focus on the
ALP encoding).
I removed this mention (it was left over LLM generated that I hadn't cut
previously). Thank you for catching it.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]