RussellSpitzer commented on code in PR #119:
URL: https://github.com/apache/parquet-testing/pull/119#discussion_r3776582788
##########
data/README.md:
##########
@@ -592,6 +593,69 @@ java -jar
parquet-cli/target/parquet-cli-1.16.0-SNAPSHOT-runtime.jar cat /home/r
{"utf8_full_truncation": "Kevin Bacon", "binary_full_truncation": "Kevin
Bacon", "utf8_partial_truncation": "πKevin Bacon", "binary_partial_truncation":
"ΓΏΓΏ\u0001\u0002", "utf8_no_truncation": "Ke", "binary_no_truncation": "Ke"}
```
+## ALP encoding
+
+`alp_extended.zstd.parquet` contains FLOAT and DOUBLE columns encoded with
+[Adaptive Lossless floating-Point
(ALP)](https://github.com/apache/parquet-format/blob/master/Encodings.md#adaptive-lossless-floating-point-alp--10)
+(`ALP = 10`).
+It was created with the code in this
[PR](https://github.com/apache/arrow/pull/49154).
+
+All columns contain the same 9032 values. The same values appear all columns
so
+decoders can bit-compare the `ALP` columns against known results stored with
+`PLAIN` encoding.
+
+
+| Column | Encoding | Rationale
/ coverage |
+|-------------------------------------|---------------------------|-----------------------------------------------------------------------------------|
+| `float_plain`, `double_plain` | `PLAIN` + zstd | In-file
reference: readers can bit-compare the ALP columns against these |
+| `float_alp_1024`, `double_alp_1024` | `ALP`, 1024-value vectors | The
default vector size of 1024 values |
+| `float_alp_4096`, `double_alp_4096` | `ALP`, 4096-value vectors | Readers
must honor `log_vector_size` from the page header rather than assume 1024 |
+| `float_alp_32`, `double_alp_32` | `ALP`, 32-value vectors | Many
vectors per page, stresses the per-vector metadata loop |
+
+
+### Data distribution (9032 rows)
+
+The "base distribution" means random values in `[-10.00, 10.00]` with exactly 2
+decimal digits (e.g. `9.43`), which are losslessly encodable by ALP (no
exceptions).
+
+The contents of the 9032 rows are as follows:
+
+| Rows | Contents
| Rationale / coverage
|
+|-----------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
+| 0β1023 | base
| Happy path: full vector, small
frame-of-reference bit width, no exceptions
|
+| 1024β2047 | base, plus: NaN at 1024, 1500 and 2047 (three distinct bit
patterns, see below), +Inf at 2000, βInf at 2001, β0.0 at 2002, subnormal
(`5e-324` double / `1e-45` float) at 2003 | NaN / Inf and sign/precision edge
values via the exception mechanism; exceptions at exact vector boundaries; NaN
payload preservation; statistics conventions (NaN excluded from min/max, Β±Inf
included, `nan_count`, β0.0/+0.0 normalization) |
+| 2048β3071 | base, plus: `3.141592653589793` at 2500
| Exactly one exception: the full-mantissa value
cannot round-trip as a decimal, and nothing else in the vector is exceptional
|
+| 3072β4095 | base, plus: `44974934523.343` at 3100 and `-1243432432.3432` at
3711
| Large-magnitude values
|
+| 4096β5119 | every 2nd value full-mantissa random, rest base
| Many (but not all) exceptions
|
+| 5120β6143 | all values full-mantissa random
| All exceptions
|
+| 6144β7167 | base but with 4 decimal digits (e.g. `3.1416`)
| Different exponent/factor than the other
vectors
|
+| 7168β8191 | constant (all `7.77`)
| `bit_width = 0` vectors
|
+| 8192β8999 | base, with nulls at every 100th row (8 nulls)
| Partial vector + null handling
|
+| 9000β9031 | random integer-valued values in `[β8e18, 8e18]`, with exactly
-8e18 at 9000 and 8e18 at 9001
| Max FOR length (64-bit) `bit_width`; (FLOAT
columns as exceptions)
|
+
+The three NaNs use distinct bit patterns so readers are checked for preserving
+non-canonical NaN payloads (ALP stores exception values bit-exactly):
+
+| Row | DOUBLE bits | FLOAT bits | Description
|
+|------|----------------------|--------------|---------------------------------|
+| 1024 | `0x7FF8000000000000` | `0x7FC00000` | Canonical quiet NaN
|
+| 1500 | `0x7FF800DEADBEEF00` | `0x7FC0DEAD` | Quiet NaN with payload
|
+| 2047 | `0xFFF8000000000001` | `0xFFC00001` | Negative quiet NaN with payload
|
+
+The file has five row groups:
+
+| Row group | Rows | Contents
|
+|-----------|-----------|---------------------------------------------------------------------|
+| 0 | 0β6143 | Base + all exception cases (6144 rows, 1546
exceptions per column) |
+| 1 | 6144β7167 | 4-decimal-digit values (different exponent/factor)
|
+| 2 | 7168β8191 | Constant `7.77` (`bit_width = 0`)
|
+| 3 | 8192β8999 | Partial trailing vector with 8 nulls
|
+| 4 | 9000β9031 | Large-magnitude values
|
+
+To check conformance of an `ALP` decoder, read each `ALP`-encoded column and
+compare the decoded values against the values from the corresponding
+`PLAIN`-encoded column. The values should be match exactly (bitwise).
Review Comment:
```suggestion
`PLAIN`-encoded column. The values should match exactly (bitwise).
```
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]