adriangb commented on issue #8608: URL: https://github.com/apache/arrow-rs/issues/8608#issuecomment-5777123166
> I don't think it is correct to write approximate distinct counts to parquet metadata in the "distinct_count' field as it doesn't conform to the spec (however useful it would be in practice) @alamb do you have a source for where in the spec it says that it has to be exact? I assumed it could be inexact becase an exact count is not very useful: even a simple query like `select count(distinct col) from t` can't be satisfied from stats if there is more than 1 file or 1 row group. Any filter also busts that. If it's only for optimizers, then an approximate count seems like it would be just as useful? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
