emecii opened a new pull request, #11193:
URL: https://github.com/apache/arrow-rs/pull/11193

   # Which issue does this PR close?
   
   Related to #11101; benchmark-only follow-up accepted in 
[#11052](https://github.com/apache/arrow-rs/pull/11052#issuecomment-5763836840).
 This covers the agreed list-index validity slice, not the full epic.
   
   # Rationale for this change
   
   #11052 corrects missing-index validity for shredded lists. Review needs 
reproducible runtime and allocation measurements of the affected paths, 
including cases where the baseline output is incorrect.
   
   # What changes are included in this PR?
   
   - A registered Criterion target with 104 deterministic cases: 64/8192 rows, 
List/ListView, unsliced/offset-three inputs, Variant/Int64 output, 
in-bounds/missing indexes, parent and explicit nulls, nested lists, 
object-to-list traversal, typed/binary fallback, and a struct-only control.
   - A separate allocation probe sharing exactly the same fixtures, with five 
samples per case, allocation/reallocation counts, requested bytes, peak 
additional live bytes, and newly retained bytes before/after output 
destruction. Timing has no tracking allocator.
   - Preflight checks against explicit expected values and an unshredded 
reference. Only the precise known #11050 baseline discrepancy is tolerated and 
counted; strict mode requires it to be fixed. Fixture setup and validation are 
outside measurement; options cloning and output destruction are included in 
timing.
   - Reproduction instructions, stable case IDs, throughput, and measurement 
boundaries. Shared input-buffer retention and private type-layout measurement 
are outside this slice. No production changes or new dependencies.
   
   # Are these changes tested?
   
   On main `1a41738b7911ebb39f172324ce8de7ef7ef2c501` plus this benchmark:
   
   - `cargo test --locked -p parquet-variant-compute --lib`: 367 passed.
   - `cargo clippy --locked -p parquet-variant-compute --all-targets 
--all-features -- -D warnings`
   - `cargo fmt --all -- --check` and `git diff --check`
   - `cargo bench --locked -p parquet-variant-compute --bench 
variant_get_list_validity -- --test`: 104 passed.
   - `cargo run --locked --release -p parquet-variant-compute --example 
variant_get_allocations`: 520 stable samples; newly retained memory released 
after each output drop.
   
   The identical suite and lockfile also ran at #11052 base 
`f9e02ba76ad11e1380b559735ac912c602b604fb` and head 
`d881317dc67e292bf0f305f2fb21b59b6a7f1458`, using separate target directories. 
All 104 cases completed; 32 baseline cases have known validity differences and 
all head cases pass strict validity checks. Strict mode rejects the baseline as 
a negative control.
   
   Local measurements: Apple M4, aarch64 macOS 26.5.2, Rust 1.98.0, optimized 
default features. Full timing runs used 30 samples, 0.2 s warm-up, 0.5 s 
measurement, 10,000 bootstrap resamples, base then head. These short sequential 
runs are exploratory; they do not establish overall performance neutrality.
   
   Selected unsliced List results (mean µs [95% confidence interval]):
   
   | Case | Base | #11052 head | Extra allocations / requested bytes |
   | --- | --- | --- | --- |
   | 64 rows, all-OOB Int64, focused repeat | 1.000 [0.996, 1.004] | 1.238 
[1.230, 1.246] | +5 / +224 |
   | 8192 rows, all in bounds, Variant | 32.314 [32.186, 32.435] | 32.392 
[32.287, 32.505] | 0 / 0 |
   | 8192 rows, mixed null/empty/value, Variant* | 54.145 [53.621, 54.626] | 
55.715 [55.322, 56.075] | +2 / +1080 |
   | 8192 rows, nested lists, Variant* | 197.688 [196.601, 199.228] | 197.784 
[196.795, 199.081] | +4 / +2160 |
   
   *Baseline validity is wrong, so these compare different output semantics. 
The 64-row typed all-OOB case has equivalent output and showed ~24% extra time 
in a focused reverse-order repeat (head then base; 50 samples, 1 s warm-up, 3 s 
measurement). This is a cost of the correctness patch for review, not a 
production change in this PR. No optimization is bundled here.
   
   # Are there any user-facing changes?
   
   No runtime/API changes. New benchmark and allocation-probe commands are 
documented in `parquet-variant-compute/benches/README.md`.
   
   # AI usage
   
   OpenAI Codex generated the benchmark fixtures, timing target, allocation 
probe, and reproduction notes; ran the validation and comparisons above; and 
drafted this description.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to