adriangb commented on issue #8608:
URL: https://github.com/apache/arrow-rs/issues/8608#issuecomment-5777123166

   > I don't think it is correct to write approximate distinct counts to 
parquet metadata in the "distinct_count' field as it doesn't conform to the 
spec (however useful it would be in practice)
   
   @alamb do you have a source for where in the spec it says that it has to be 
exact?
   
   I assumed it could be inexact becase an exact count is not very useful: even 
a simple query like `select count(distinct col) from t` can't be satisfied from 
stats if there is more than 1 file or 1 row group. Any filter also busts that. 
If it's only for optimizers, then an approximate count seems like it would be 
just as useful?


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to