etseidl commented on code in PR #258: URL: https://github.com/apache/parquet-format/pull/258#discussion_r1629834802
########## CONTRIBUTING.md: ########## @@ -29,3 +29,99 @@ Recommendations and requirements for how to best contribute to Parquet. We striv ### License By contributing your code, you agree to license your contribution under the terms of the APLv2: https://github.com/apache/parquet-format/blob/master/LICENSE + +### Additions/Changes to the Format + +Note: This section applies to actual functional changes to the +specification. Fixing typos, grammar, and clarifying concepts +that would not change the semantics of the specification can +be done as long a comitter feels comfortable to merge them. When +in doubt starting a discussion on the dev mailing list is +encouraged. + +The general steps for adding features to the format are as follows: + +1. Discuss changes on on the developer mailing list ([email protected]). Often times it is helpful to link to a draft pull request to make the discussion concrete. This step is complete when there lazy consensus. + +2. Once a change has lazy consensus two implementations of the feature +demonstrating interopability must also be provided. One implementation MUST be [parquet-java](http://github.com/apache/parquet-java). It is preferred that the second implementation be [parquet-cpp](https://github.com/apache/arrow) or [parquet-rs](https://github.com/apache/arrow-rs), however at the discretion of the PMC any +open source Parquet implementation may be acceptable. Implementations +whose contributors actively +participate in the community (e.g. keep their feature matrix +up-to-date on parquet-site) are more likely to be considered. + +Unless otherwise discussed, it is expected the implementations will +develop from the main branch (i.e. backporting is not expected). + +In some cases in addition to library level implementations it is +expected the changes will be justified via integration into a +processing engine to show their viability. Review Comment: IIUC, the intent here is to demonstrate some real world benefit to justify the cost. For instance, for the recent size statistics addition we probably should have required some benchmark numbers showing a clear benefit to justify the cost of creating and storing them. My GPU workflow was made much faster due to the inclusion of the unencoded byte array size information and could have provided hard numbers demonstrating that. Similarly, it would have been nice to have a pushdown example that was made possible because of the histograms. Simply adding the feature and showing that two implementations produced them in a compatible way was probably insufficient. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
