thisisnic commented on issue #34319:
URL: https://github.com/apache/arrow/issues/34319#issuecomment-5822219231

   Note: this comment was written by Claude Code (an AI assistant) at my 
request, after I asked it to check whether this issue still applies. I have 
reviewed the findings but the measurements and analysis are its own.
   
   I re-ran this on current main (1f789ffd70) to see whether it still applies. 
The short answer is that the `col_select` regression described in the original 
report is gone, but the reproducer is still slow for a different reason.
   
   With 16000 plain integer columns and no attributes, `read_parquet(file, 
col_select = 1:200)` takes about 0.15 s and reading all columns takes about 
0.27 s. So the `eval_select` path that was suspected above is no longer a 
problem.
   
   The reproducer in this issue attaches `levels` and `text` attributes to 
every column, and that is where the time now goes. With 16000 such columns and 
10 rows:
   
   | | elapsed |
   |---|---|
   | `as.data.frame(reader$ReadTable(1:200))` | 1.28 s |
   | `read_parquet(file, col_select = 1:200)` | 1.34 s |
   | `tab$metadata$r` alone, on the 200 column table | 1.24 s |
   | `read_parquet(file)`, all columns | 10.5 s |
   
   Two separate costs, both in R metadata handling and both scaling with the 
number of columns in the file rather than the number selected:
   
   1. Unserialising and sanitising the R metadata costs about 1.2 s per read 
regardless of `col_select`, because the table still carries metadata for all 
16000 columns and `safe_r_metadata()` walks all of it in R: 
https://github.com/apache/arrow/blob/1f789ffd70f96907698e93ef99f4e034caf0788b/r/R/metadata.R#L73.
 This was added in #41969, so it postdates the original report and is why the 
`ParquetFileReader` workaround above no longer helps much for data like this.
   
   2. Applying the per-column metadata costs a further 9 s when reading all 
columns, because the loop assigns into the data frame with `x[[name]] <-` on 
every iteration, copying the whole 16000 column frame each time: 
https://github.com/apache/arrow/blob/1f789ffd70f96907698e93ef99f4e034caf0788b/r/R/metadata.R#L178-L182.
 #48104 added an early exit for the case where no column has metadata, so it 
does not help here.
   
   I think this should stay open, but the title no longer describes the 
problem. Suggest retitling to something like "[R] Reading wide files with 
per-column R metadata is slow" and treating the two costs above as the work 
items. The second one is a small fix: collect the modified columns and assign 
once after the loop.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to