adriangb opened a new issue, #10511:
URL: https://github.com/apache/arrow-rs/issues/10511

   **Is your feature request related to a problem or challenge? Please describe 
what you are trying to do.**
   
   Follow-up to #10505 (see the "Notes" section there), split out so it can be 
evaluated, and possibly rejected, on its own data.
   
   The byte-budget sub-batching that bounds data page sizes (#9972) measures 
raw payload bytes: `count_values_within_byte_budget*` sums each value's 
plain-encoded length and cuts a mini-batch when the sum crosses the page byte 
limit. For `PLAIN` and `DELTA_LENGTH_BYTE_ARRAY` that is exact. For 
`DELTA_BYTE_ARRAY` it can overestimate arbitrarily: a column of large values 
sharing long prefixes encodes to a fraction of its raw size, but the chunker 
still slices it by pre-dedup size, down to one-value mini-batches for values 
above the limit.
   
   **How much this actually costs**
   
   Since #10505, mini-batch sizing no longer affects file size or page layout: 
page-cut decisions consult the encoder's real encoded estimate, so a 
well-compressing column accumulates a full page regardless of how finely the 
input was sub-batched. What remains attributable to the raw-byte budget is 
per-mini-batch dispatch overhead (`write_mini_batch`, level encoding, the 
post-write page check).
   
   Measured with the benchmark from #BENCH_PR_NUMBER (128 rows x 2 MiB, delta 
vs `PLAIN` on identical data, Apple M-series, on top of #10505):
   
   | group | `plain` | `delta_byte_array` |
   | --- | --- | --- |
   | `large_string_shared_prefix` | 58.8 ms (4.2 GiB/s) | 91.2 ms (2.7 GiB/s) |
   | `large_string_distinct` | 39.8 ms (6.3 GiB/s) | 27.5 ms (9.1 GiB/s) |
   
   The shared-prefix delta deficit is not sub-batching overhead: it is the 
encoder's inherent prefix comparison (a ~2 MiB memcmp per value pair) plus 
suffix bookkeeping, work an encoded-size-aware budget would not remove. The 
dispatch overhead the budget is responsible for is one `write_mini_batch` per 
value here, which is noise at these value sizes; at smaller value sizes the 
budget yields proportionally larger mini-batches, so the dispatch count stays 
low there too.
   
   **Describe the solution you'd like**
   
   Possibly nothing: on this evidence the raw-byte budget costs throughput only 
through mini-batch dispatch, and that cost is not measurable at the scales 
benchmarked. Filing so the question raised in #10505's notes has a home and 
data, rather than as a commitment to build it. If someone has a workload where 
sub-batch dispatch shows up in profiles, this is the place for it.
   
   If it does prove worth doing, the shape of the fix and its obstacles, from 
#10505's analysis:
   
   - The counting functions are stateless associated functions; the prefix 
length of a chunk's first value depends on the encoder's live `previous` value, 
so encoder state would need to be threaded through 
`ByteBudgetChunker::pick_sub_batch_size` and both trait impls (the trait is 
crate-private, so no API concern).
   - Computing prefix lengths in the counter duplicates the byte comparisons 
the delta encoder does again at write time, in a path designed to short-circuit 
cheaply.
   - Any overestimate of achievable dedup re-opens the page-overshoot hole 
#9972 closed, so the estimate must stay conservative.
   
   A cheaper alternative to prediction is feedback: size the next mini-batch 
from the page's remaining encoded budget (`limit - 
estimated_data_page_size()`), which uses the encoder's own accurate estimate 
and needs no prefix prediction, at the cost of trailing one mini-batch behind 
reality.
   
   **Describe alternatives you've considered**
   
   Close as not-planned once #10505 merges, on the measurements above.
   
   **Additional context**
   
   #9972 introduced the byte budget; #10489 / #10505 are the `DELTA_BYTE_ARRAY` 
page-split regression and fix this follows from; #BENCH_PR_NUMBER adds the 
benchmark used for the measurements.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to