This is an automated email from the ASF dual-hosted git repository.

thisisnic pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/arrow.git


The following commit(s) were added to refs/heads/main by this push:
     new 75387422926 GH-30800: [Python][Docs] Document partition fields with 
explicit dataset schemas (#50352)
75387422926 is described below

commit 7538742292681feaa3ad5f008c18e105b16ec54a
Author: Kalyanam Dewri <[email protected]>
AuthorDate: Wed Sep 30 04:22:36 2026 -0700

    GH-30800: [Python][Docs] Document partition fields with explicit dataset 
schemas (#50352)
    
    Summary:
    - Document that explicit dataset schemas for partitioned datasets must 
include partition fields used in filters or projections.
    - Add a doctest-backed example showing a Hive partition field used with 
`count_rows`.
    
    Validation:
    - `/tmp/arrow-doccheck-30800/bin/python -m pytest --doctest-glob="*.rst" 
docs/source/python/dataset.rst -q` — passed, `1 passed`
    - `git diff --check HEAD~1 HEAD` — passed
    
    * GitHub Issue: #30800
    
    Authored-by: kalyanamdewri <[email protected]>
    Signed-off-by: Nic Crane <[email protected]>
---
 docs/source/python/dataset.rst | 18 ++++++++++++++++++
 1 file changed, 18 insertions(+)

diff --git a/docs/source/python/dataset.rst b/docs/source/python/dataset.rst
index 4e18ea0a51c..0c51458c8e6 100644
--- a/docs/source/python/dataset.rst
+++ b/docs/source/python/dataset.rst
@@ -374,6 +374,24 @@ altogether if they do not match the filter:
     3  8  0.313068  1    b
     4  9 -0.854096  2    b
 
+When passing an explicit ``schema`` to :func:`dataset`, include the partition
+fields in the schema if they are used in filters or projections. The partition
+fields are not stored in the physical files, so they need to be present in the
+dataset schema when schema inference is bypassed:
+
+.. code-block:: python
+
+    >>> schema = pa.schema([
+    ...     ("a", pa.int64()),
+    ...     ("b", pa.float64()),
+    ...     ("c", pa.int64()),
+    ...     ("part", pa.string()),
+    ... ])
+    >>> dataset = ds.dataset("parquet_dataset_partitioned", format="parquet",
+    ...                      partitioning="hive", schema=schema)
+    >>> dataset.count_rows(filter=ds.field("part") == "b")
+    5
+
 
 Different partitioning schemes
 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

Reply via email to