thisisnic commented on issue #32795:
URL: https://github.com/apache/arrow/issues/32795#issuecomment-5822339781
This comment was written by Claude Code, an AI assistant, which ran the
investigation below under my direction. I've reviewed the results before
posting.
I revisited this with Arrow 25.0.0. The original MinIO endpoint is no longer
reachable, so I reproduced locally with a single large parquet file (40M and
80M rows, 8 columns, 100k-row row groups), running open_dataset() into
write_dataset() and sampling process RSS every 0.5s.
| Run | File on disk | Peak RSS | Pool peak |
|---|---|---|---|
| R, defaults | 1.4 GB | 3.1 GB | 2.3 GB |
| R, defaults | 2.8 GB | 4.5 GB | 3.7 GB |
| pyarrow 25.0.1, defaults | 2.8 GB | 3.8 GB | 3.7 GB |
| R, pre_buffer = FALSE | 2.8 GB | 1.9 GB | 0.98 GB |
| pyarrow, pre_buffer=False | 2.8 GB | 1.3 GB | 0.99 GB |
What I found:
- Peak memory grows with file size, and grows steadily through the run
rather than plateauing, which matches the original report.
- R and pyarrow now have identical memory pool peaks, so the R-specific
retention seen in 2022 (R holding batches Python released) no longer happens.
Allocator choice (mimalloc, jemalloc, system) makes no difference, and nothing
is retained after write_dataset() returns.
- The remaining growth is parquet pre-buffering. It is on by default and,
for a single large file, caches the raw bytes of every row group without
evicting the ones already consumed. Disabling it flattens the curve with no
slowdown on local disk:
open_dataset(path, format = FileFormat$create("parquet", pre_buffer =
FALSE))
- This is the same mechanism as #39808, and #49855 is the C++ fix (evict
pre-buffered row-group bytes after decode).
The 9.0.0 doubling was most likely the readahead change Weston mentioned,
amplified by this retention. The dataset writer backpressure fix in 16.0.0
(#40224) addressed the writer side separately.
Proposal: close this as a duplicate of #39808 once #49855 lands, and link
#33366 which looks like the same cause.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]