neilconway opened a new pull request, #11360:
URL: https://github.com/apache/arrow-rs/pull/11360

   …le values
   
   `take` on a List, LargeList or Map estimated the output's child capacity as 
the average row length times the number of rows taken. It worked out the 
average from the whole child array, including values outside a slice, so taking 
rows from a small slice of a large array could reserve far more memory than 
needed, and the result kept that memory. For example, taking 100 rows from a 
two-row slice of a 1000-row array of 10-element lists used about 4 MB instead 
of 8 KB.
   
   Work out the average from the child values that the rows use.
   
   # Which issue does this PR close?
   
   - N/A
   
   # Rationale for this change
   
   `take` on a list or map uses the average input row length as part of its 
estimate for sizing its output buffer. However, it calculated the average row 
length starting from raw size of the input values buffer, not the visible span. 
This could result in computing a very large "average row length" for sliced 
lists and maps, and therefore over-allocating output capacity.
   
   # What changes are included in this PR?
   
   * Use the visible span to compute average row length
   * Refactor capacity calculations to avoid duplication
   * Add unit test
   
   # Are these changes tested?
   
   Yes; existing tests pass, new test added.
   
   # Are there any significant user-facing changes?
   
   No.
   
   # AI usage
   
   Developed and revised with Claude Code (Opus 5.5), reviewed with Codex.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to