This is an automated email from the ASF dual-hosted git repository.
thisisnic pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/arrow.git
The following commit(s) were added to refs/heads/main by this push:
new 75387422926 GH-30800: [Python][Docs] Document partition fields with
explicit dataset schemas (#50352)
75387422926 is described below
commit 7538742292681feaa3ad5f008c18e105b16ec54a
Author: Kalyanam Dewri <[email protected]>
AuthorDate: Wed Sep 30 04:22:36 2026 -0700
GH-30800: [Python][Docs] Document partition fields with explicit dataset
schemas (#50352)
Summary:
- Document that explicit dataset schemas for partitioned datasets must
include partition fields used in filters or projections.
- Add a doctest-backed example showing a Hive partition field used with
`count_rows`.
Validation:
- `/tmp/arrow-doccheck-30800/bin/python -m pytest --doctest-glob="*.rst"
docs/source/python/dataset.rst -q` — passed, `1 passed`
- `git diff --check HEAD~1 HEAD` — passed
* GitHub Issue: #30800
Authored-by: kalyanamdewri <[email protected]>
Signed-off-by: Nic Crane <[email protected]>
---
docs/source/python/dataset.rst | 18 ++++++++++++++++++
1 file changed, 18 insertions(+)
diff --git a/docs/source/python/dataset.rst b/docs/source/python/dataset.rst
index 4e18ea0a51c..0c51458c8e6 100644
--- a/docs/source/python/dataset.rst
+++ b/docs/source/python/dataset.rst
@@ -374,6 +374,24 @@ altogether if they do not match the filter:
3 8 0.313068 1 b
4 9 -0.854096 2 b
+When passing an explicit ``schema`` to :func:`dataset`, include the partition
+fields in the schema if they are used in filters or projections. The partition
+fields are not stored in the physical files, so they need to be present in the
+dataset schema when schema inference is bypassed:
+
+.. code-block:: python
+
+ >>> schema = pa.schema([
+ ... ("a", pa.int64()),
+ ... ("b", pa.float64()),
+ ... ("c", pa.int64()),
+ ... ("part", pa.string()),
+ ... ])
+ >>> dataset = ds.dataset("parquet_dataset_partitioned", format="parquet",
+ ... partitioning="hive", schema=schema)
+ >>> dataset.count_rows(filter=ds.field("part") == "b")
+ 5
+
Different partitioning schemes
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~