etseidl commented on PR #11297:
URL: https://github.com/apache/arrow-rs/pull/11297#issuecomment-5938177071

   > I did have a question which is what is the usecase when we want to make 
`ParquetMetaData` easier to clone? Most uses already have 
`Arc<ParquetMetaData>` I think
   
   I was trying to address point 1 in 
https://github.com/apache/datafusion/issues/24288#issuecomment-5269340232. If 
the `Arc<ParquetMetaData>` is already shared (i.e. refcount > 1), then you 
can't un-Arc it to attach a new page index provider. In that case, 
`unwrap_or_clone` will clone the entire row group and file metadata, which can 
be costly. If `ParquetMetaData` is a wrapper around shared components, cloning 
becomes much cheaper and we can have easy throw-away `ParquetMetaData` 
instances to throw around.
   
   > Is the idea that one of the subfields like the FileMetadata or the 
row_group_metadata will be modified in one instance of the ParquetMetadata but 
not others 🤔
   
   That's another possibility. If the metadata decoders are modified to return 
parts rather than the assembled whole, we could at some future point cache the 
individual parts and mutate them to fit a particular query. Wrap the custom 
parts in a new `ParquetMetaData` and send them on their way.
   
   A distant fever dream of mine is to have each component (`RowGroupMetaData`, 
`ColumnChunkMetaData`, `ColumnIndexMetaData`, etc) individually wrapped in an 
`Arc` and kept in a store. Custom metadata can then be produced by reaching 
into the store and cloning the necessary bits. The store would be populated as 
needed, so a batch of queries against a small part of a wide table would never 
decode the full metadata. This would be possible if we implement the footer 
skip index, or wait to see what the new modular footer enables. (I really need 
to catch up on that work)


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to