thisisnic commented on issue #32795:
URL: https://github.com/apache/arrow/issues/32795#issuecomment-5822339781

   This comment was written by Claude Code, an AI assistant, which ran the 
investigation below under my direction. I've reviewed the results before 
posting.
   
   I revisited this with Arrow 25.0.0. The original MinIO endpoint is no longer 
reachable, so I reproduced locally with a single large parquet file (40M and 
80M rows, 8 columns, 100k-row row groups), running open_dataset() into 
write_dataset() and sampling process RSS every 0.5s.
   
   | Run | File on disk | Peak RSS | Pool peak |
   |---|---|---|---|
   | R, defaults | 1.4 GB | 3.1 GB | 2.3 GB |
   | R, defaults | 2.8 GB | 4.5 GB | 3.7 GB |
   | pyarrow 25.0.1, defaults | 2.8 GB | 3.8 GB | 3.7 GB |
   | R, pre_buffer = FALSE | 2.8 GB | 1.9 GB | 0.98 GB |
   | pyarrow, pre_buffer=False | 2.8 GB | 1.3 GB | 0.99 GB |
   
   What I found:
   
   - Peak memory grows with file size, and grows steadily through the run 
rather than plateauing, which matches the original report.
   - R and pyarrow now have identical memory pool peaks, so the R-specific 
retention seen in 2022 (R holding batches Python released) no longer happens. 
Allocator choice (mimalloc, jemalloc, system) makes no difference, and nothing 
is retained after write_dataset() returns.
   - The remaining growth is parquet pre-buffering. It is on by default and, 
for a single large file, caches the raw bytes of every row group without 
evicting the ones already consumed. Disabling it flattens the curve with no 
slowdown on local disk:
   
     open_dataset(path, format = FileFormat$create("parquet", pre_buffer = 
FALSE))
   
   - This is the same mechanism as #39808, and #49855 is the C++ fix (evict 
pre-buffered row-group bytes after decode).
   
   The 9.0.0 doubling was most likely the readahead change Weston mentioned, 
amplified by this retention. The dataset writer backpressure fix in 16.0.0 
(#40224) addressed the writer side separately.
   
   Proposal: close this as a duplicate of #39808 once #49855 lands, and link 
#33366 which looks like the same cause.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to