dannyshafer opened a new pull request, #51644: URL: https://github.com/apache/arrow/pull/51644
### Rationale for this change The Python Parquet reader docstrings (`ParquetFile`, `ParquetDataset`, `read_table` and `read_pandas`) each repeated the descriptions of common options such as `thrift_string_size_limit`, and the copies had drifted slightly apart. Closes #51265. ### What changes are included in this PR? - The shared parameter descriptions now live in module-level snippets next to the existing `_read_docstring_common`: `_read_docstring_file_options` (read_dictionary through buffer_size), `_read_docstring_partitioning`, `_read_docstring_filesystem`, `_read_docstring_reader_options` (pre_buffer through schema_depth_limit) and `_read_docstring_checksum_options` (page_checksum_verification, arrow_extensions_enabled). `_read_docstring_common` is kept, built from the first two. - `ParquetFile.__doc__` becomes an f-string built from these snippets, the same way `ParquetDataset.__doc__` already is. That is why its body is dedented in the diff. - `ParquetDataset.__doc__`, `read_table.__doc__` and `read_pandas.__doc__` use the same snippets. - Where copies differed, the most complete wording was kept. For example, `read_table` now gets the full `pre_buffer` text (S3, GCS and the memory-usage note), `ParquetFile` gets the full `read_dictionary` description (I checked that `ParquetFile(..., read_dictionary=["s.x"])` does accept nested column paths), and `decryption_properties` mentions both Parquet Modular Encryption and `CryptoFactory.file_decryption_properties()`. ### Are these changes tested? This is a docstring-only change. I rendered the four docstrings before and after and diffed them: apart from the wording unifications above, the content is identical. I also ran the `ParquetFile` docstring examples with doctest and ran `flake8 --config python/setup.cfg` and `autopep8 --diff` on the file (both clean). ### Are there any user-facing changes? Only the rendered docstrings. They are now consistent across the four readers. ### Was AI used for this PR? In accordance to the [AI generation guidelines](https://arrow.apache.org/docs/dev/developers/overview.html#ai-generated-code), please disclose below whether and how AI was used in this PR. This PR was prepared with Claude Code, which wrote the refactor, compared the rendered docstrings before and after, and drafted this description. **PR code and description written by:** - [ ] Human - [x] AI **Reviewed before submission by:** - [ ] Human - [x] AI - [ ] Not reviewed <!-- oss-routine --> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
