This is an automated email from the ASF dual-hosted git repository.
pitrou pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/arrow.git
The following commit(s) were added to refs/heads/main by this push:
new bda98c4d8c GH-37476: [C++][Python] Preserve unsigned dictionary index
types when building from values (#50475)
bda98c4d8c is described below
commit bda98c4d8c4dbf52c6b52186e07a7bb0578c0ac4
Author: Max Gorbuk <[email protected]>
AuthorDate: Tue Jul 28 20:01:45 2026 +0300
GH-37476: [C++][Python] Preserve unsigned dictionary index types when
building from values (#50475)
### Rationale for this change
An unsigned dictionary index type is silently replaced by the signed one of
the same width:
```python
>>> pa.array(["a", "b"], type=pa.dictionary(pa.uint32(), pa.string())).type
dictionary<values=string, indices=int32, ordered=0>
```
It also makes `pa.chunked_array(values, dict_type)` fail with `Array chunks
must all be same type`, since the chunk built internally comes back a different
type than requested.
`DictionaryBuilderCase::CreateFor()` reduces the requested index type to
its byte width and hands that to `AdaptiveIntBuilder`, so everything but the
width is dropped and the type is rebuilt from what the indices builder reports,
which is always signed.
@ jorisvandenbossche suggested `AdaptiveUIntBuilder` for unsigned types.
That returns a different builder class, which `util/converter.h`,
`json/from_string.cc` and the R binding all cast to `DictionaryBuilder<T>`, and
`TestDictionaryUnifier.ChunkedArrayNestedDict` hits a `checked_pointer_cast`
DCHECK. So this keeps one builder class and just reports the requested
signedness. Indices are non-negative and same-width signed and unsigned types
have the same layout, so it is value-preserv [...]
### What changes are included in this PR?
`DictionaryBuilderBase` and its `NullType` specialization record the
requested signedness and map the signed index type to the unsigned one of the
same width in `type()`, `FinishInternal()` and `FinishDelta()`; `CreateFor()`
passes it through. The mapping helper lives in `builder.cc`, not the header.
`util/converter.h`, `python_to_arrow.cc` and `r_to_arrow.cpp` are untouched, so
R is fixed rather than broken.
uint64 indices also convert to pandas now, mapped to int64 codes. That is
safe because the indices are bounds-checked below the dictionary length, so
they never reach the int64 limit; `WriteIndicesUniform` handles uint64 like it
already handles uint32.
### Are these changes tested?
Six tests in `array_dict_test.cc` cover the distinct paths: type and finish
preservation across signed and unsigned, the width adapting while staying
unsigned, the `NullType` builder, `FinishDelta`, the supplied-dictionary
constructor and the exact-index builder. Three parametrized tests in
`test_array.py` cover the eight index types, the width adapting and the
uint64-to-pandas conversion, and `test_dictionary_with_pandas` now checks that
conversion instead of the old error.
A `uint8` grows to `uint16` once more than 128 distinct values are added
and stays unsigned; the underlying builder is signed, so it widens after 128
rather than the 256 a real uint8 could hold. `arrow-array-test`,
`arrow-ipc-read-write-test`, `arrow-compute-scalar-cast-test`,
`arrow-c-bridge-test` and the pyarrow `test_array` and `test_pandas` suites
pass.
### Are there any user-facing changes?
`pa.array` and `pa.chunked_array` return the requested unsigned index type.
`dictionary(uint64(), ...)` produces real uint64 indices, and converting such
an array to pandas now works where it used to raise. `take`, `filter`, `cast`,
`dictionary_decode`, IPC and scalar access all keep the index type.
* GitHub Issue: #37476
Authored-by: mkzung <[email protected]>
Signed-off-by: Antoine Pitrou <[email protected]>
---
cpp/src/arrow/array/array_dict_test.cc | 117 +++++++++++++++++++++
cpp/src/arrow/array/builder_dict.h | 48 ++++++---
cpp/src/arrow/builder.cc | 43 +++++++-
python/pyarrow/src/arrow/python/arrow_to_pandas.cc | 14 +--
python/pyarrow/tests/test_array.py | 39 +++++++
python/pyarrow/tests/test_pandas.py | 13 +--
6 files changed, 242 insertions(+), 32 deletions(-)
diff --git a/cpp/src/arrow/array/array_dict_test.cc
b/cpp/src/arrow/array/array_dict_test.cc
index 5f4335bcbc..23335ebb00 100644
--- a/cpp/src/arrow/array/array_dict_test.cc
+++ b/cpp/src/arrow/array/array_dict_test.cc
@@ -1040,6 +1040,123 @@ TYPED_TEST(TestDictionaryBuilderIndexByteWidth,
MakeBuilder) {
AssertIndexByteWidth<TypeParam, NullType>();
}
+// ----------------------------------------------------------------------
+// GH-37476: the requested dictionary index type's signedness must be
preserved.
+
+TEST(TestDictionaryBuilderIndexType, PreservesRequestedIndexType) {
+ // Both the builder's reported type and the finished array must carry the
requested
+ // index type, for signed and unsigned widths alike.
+ for (auto index_type :
+ {int8(), int16(), int32(), int64(), uint8(), uint16(), uint32(),
uint64()}) {
+ ARROW_SCOPED_TRACE("index_type = ", index_type->ToString());
+ auto dict_type = dictionary(index_type, utf8());
+ ASSERT_OK_AND_ASSIGN(auto builder, MakeBuilder(dict_type));
+ AssertTypeEqual(*index_type,
+ *checked_cast<const
DictionaryType&>(*builder->type()).index_type());
+
+ auto& dict_builder =
checked_cast<DictionaryBuilder<StringType>&>(*builder);
+ ASSERT_OK(dict_builder.Append("a"));
+ ASSERT_OK(dict_builder.Append("b"));
+ ASSERT_OK(dict_builder.AppendNull());
+ ASSERT_OK(dict_builder.Append("a"));
+ ASSERT_OK_AND_ASSIGN(auto result, dict_builder.Finish());
+ ASSERT_OK(result->ValidateFull());
+
+ auto ex_dict = ArrayFromJSON(utf8(), R"(["a", "b"])");
+ auto ex_indices = ArrayFromJSON(index_type, "[0, 1, null, 0]");
+ DictionaryArray expected(dict_type, ex_indices, ex_dict);
+ AssertTypeEqual(*dict_type, *result->type());
+ AssertArraysEqual(expected, *result);
+ }
+}
+
+TEST(TestDictionaryBuilderIndexType, WidthAdaptsWhenUnsigned) {
+ // The width stays adaptive, as it does for signed indices. The underlying
builder is
+ // signed, so it widens after 128 distinct values rather than the 256 a
uint8 could
+ // hold, but the widened type stays unsigned rather than falling back to a
signed type.
+ auto dict_type = dictionary(uint8(), utf8());
+ ASSERT_OK_AND_ASSIGN(auto boxed_builder, MakeBuilder(dict_type));
+ auto& builder = checked_cast<DictionaryBuilder<StringType>&>(*boxed_builder);
+
+ for (int i = 0; i < 200; ++i) {
+ ASSERT_OK(builder.Append(std::to_string(i)));
+ }
+
+ ASSERT_OK_AND_ASSIGN(auto result, builder.Finish());
+ ASSERT_OK(result->ValidateFull());
+ AssertTypeEqual(*uint16(),
+ *checked_cast<const
DictionaryType&>(*result->type()).index_type());
+}
+
+TEST(TestDictionaryBuilderIndexType, NullValueTypePreservesUnsignedIndexType) {
+ // The NullType value builder is a separate specialization with its own
type() and
+ // FinishInternal, so it needs its own guard.
+ auto dict_type = dictionary(uint32(), null());
+ ASSERT_OK_AND_ASSIGN(auto boxed_builder, MakeBuilder(dict_type));
+ auto& builder = checked_cast<DictionaryBuilder<NullType>&>(*boxed_builder);
+ AssertTypeEqual(*uint32(),
+ *checked_cast<const
DictionaryType&>(*builder.type()).index_type());
+
+ ASSERT_OK(builder.AppendNull());
+ ASSERT_OK_AND_ASSIGN(auto result, builder.Finish());
+ ASSERT_OK(result->ValidateFull());
+ AssertTypeEqual(*uint32(),
+ *checked_cast<const
DictionaryType&>(*result->type()).index_type());
+}
+
+TEST(TestDictionaryBuilderIndexType, FinishDeltaPreservesUnsignedIndexType) {
+ // FinishDelta is a distinct path from Finish and must carry the requested
index type
+ // too, not the signed type the adaptive builder produces internally.
+ auto dict_type = dictionary(uint32(), utf8());
+ ASSERT_OK_AND_ASSIGN(auto boxed_builder, MakeBuilder(dict_type));
+ auto& builder = checked_cast<DictionaryBuilder<StringType>&>(*boxed_builder);
+
+ ASSERT_OK(builder.Append("a"));
+ ASSERT_OK(builder.Append("b"));
+
+ std::shared_ptr<Array> result_indices, result_delta;
+ ASSERT_OK(builder.FinishDelta(&result_indices, &result_delta));
+ ASSERT_OK(result_indices->ValidateFull());
+ AssertTypeEqual(*uint32(), *result_indices->type());
+ AssertArraysEqual(*ArrayFromJSON(uint32(), "[0, 1]"), *result_indices);
+}
+
+TEST(TestDictionaryBuilderIndexType,
SuppliedDictionaryPreservesUnsignedIndexType) {
+ // The supplied-dictionary constructor starts the adaptive builder at its
default width
+ // rather than the requested one, so a requested uint32 reports uint8. The
width is not
+ // honoured on this path (a signed request behaves the same way), but the
signedness
+ // must survive regardless: uint8, not int8.
+ auto dict_values = ArrayFromJSON(utf8(), R"(["a", "b"])");
+ ASSERT_OK_AND_ASSIGN(auto builder,
+ MakeDictionaryBuilder(dictionary(uint32(), utf8()),
dict_values));
+ AssertTypeEqual(*uint8(),
+ *checked_cast<const
DictionaryType&>(*builder->type()).index_type());
+
+ auto& dict_builder = checked_cast<DictionaryBuilder<StringType>&>(*builder);
+ ASSERT_OK(dict_builder.Append("a"));
+ ASSERT_OK(dict_builder.Append("b"));
+ ASSERT_OK_AND_ASSIGN(auto result, dict_builder.Finish());
+ ASSERT_OK(result->ValidateFull());
+ AssertTypeEqual(*uint8(),
+ *checked_cast<const
DictionaryType&>(*result->type()).index_type());
+}
+
+TEST(TestDictionaryBuilderIndexType,
ExactIndexBuilderPreservesUnsignedIndexType) {
+ // MakeBuilderExactIndex is a separate builder that honours unsigned index
types on its
+ // own; guard that it keeps doing so.
+ auto dict_type = dictionary(uint16(), utf8());
+ std::unique_ptr<ArrayBuilder> boxed_builder;
+ ASSERT_OK(MakeBuilderExactIndex(default_memory_pool(), dict_type,
&boxed_builder));
+ AssertTypeEqual(
+ *uint16(),
+ *checked_cast<const
DictionaryType&>(*boxed_builder->type()).index_type());
+
+ ASSERT_OK_AND_ASSIGN(auto result, boxed_builder->Finish());
+ ASSERT_OK(result->ValidateFull());
+ AssertTypeEqual(*uint16(),
+ *checked_cast<const
DictionaryType&>(*result->type()).index_type());
+}
+
// ----------------------------------------------------------------------
// DictionaryArray tests
diff --git a/cpp/src/arrow/array/builder_dict.h
b/cpp/src/arrow/array/builder_dict.h
index 269cdeee64..4dd4e6ea44 100644
--- a/cpp/src/arrow/array/builder_dict.h
+++ b/cpp/src/arrow/array/builder_dict.h
@@ -134,6 +134,12 @@ class ARROW_EXPORT DictionaryMemoTable {
namespace internal {
+/// \brief Return the unsigned integer type of the same width when an unsigned
+/// dictionary index type was requested. Defined in builder.cc; see there for
why
+/// this is value-preserving and why the width widens on the signed threshold.
+ARROW_EXPORT std::shared_ptr<DataType> MaybeUnsignedIndexType(
+ const std::shared_ptr<DataType>& index_type, bool use_unsigned_index);
+
/// \brief Array builder for created encoded DictionaryArray from
/// dense array
///
@@ -154,14 +160,16 @@ class DictionaryBuilderBase : public ArrayBuilder {
const std::shared_ptr<DataType>&>
value_type,
MemoryPool* pool = default_memory_pool(),
- int64_t alignment = kDefaultBufferAlignment, bool
ordered = false)
+ int64_t alignment = kDefaultBufferAlignment, bool
ordered = false,
+ bool use_unsigned_index = false)
: ArrayBuilder(pool, alignment),
memo_table_(new internal::DictionaryMemoTable(pool, value_type)),
delta_offset_(0),
byte_width_(-1),
indices_builder_(start_int_size, pool, alignment),
value_type_(value_type),
- ordered_(ordered) {}
+ ordered_(ordered),
+ use_unsigned_index_(use_unsigned_index) {}
template <typename T1 = T>
explicit DictionaryBuilderBase(
@@ -199,14 +207,16 @@ class DictionaryBuilderBase : public ArrayBuilder {
const std::shared_ptr<DataType>&>
value_type,
MemoryPool* pool = default_memory_pool(),
- int64_t alignment = kDefaultBufferAlignment, bool
ordered = false)
+ int64_t alignment = kDefaultBufferAlignment, bool
ordered = false,
+ bool use_unsigned_index = false)
: ArrayBuilder(pool, alignment),
memo_table_(new internal::DictionaryMemoTable(pool, value_type)),
delta_offset_(0),
byte_width_(static_cast<const T1&>(*value_type).byte_width()),
indices_builder_(start_int_size, pool, alignment),
value_type_(value_type),
- ordered_(ordered) {}
+ ordered_(ordered),
+ use_unsigned_index_(use_unsigned_index) {}
template <typename T1 = T>
explicit DictionaryBuilderBase(
@@ -244,14 +254,15 @@ class DictionaryBuilderBase : public ArrayBuilder {
explicit DictionaryBuilderBase(const std::shared_ptr<Array>& dictionary,
MemoryPool* pool = default_memory_pool(),
int64_t alignment = kDefaultBufferAlignment,
- bool ordered = false)
+ bool ordered = false, bool use_unsigned_index
= false)
: ArrayBuilder(pool, alignment),
memo_table_(new internal::DictionaryMemoTable(pool, dictionary)),
delta_offset_(0),
byte_width_(-1),
indices_builder_(pool, alignment),
value_type_(dictionary->type()),
- ordered_(ordered) {}
+ ordered_(ordered),
+ use_unsigned_index_(use_unsigned_index) {}
~DictionaryBuilderBase() override = default;
@@ -486,6 +497,7 @@ class DictionaryBuilderBase : public ArrayBuilder {
std::shared_ptr<ArrayData> indices_data;
std::shared_ptr<ArrayData> delta_data;
ARROW_RETURN_NOT_OK(FinishWithDictOffset(delta_offset_, &indices_data,
&delta_data));
+ indices_data->type = MaybeUnsignedIndexType(indices_data->type,
use_unsigned_index_);
*out_indices = MakeArray(indices_data);
*out_delta = MakeArray(delta_data);
return Status::OK();
@@ -498,7 +510,9 @@ class DictionaryBuilderBase : public ArrayBuilder {
Status Finish(std::shared_ptr<DictionaryArray>* out) { return
FinishTyped(out); }
std::shared_ptr<DataType> type() const override {
- return ::arrow::dictionary(indices_builder_.type(), value_type_, ordered_);
+ return ::arrow::dictionary(
+ MaybeUnsignedIndexType(indices_builder_.type(), use_unsigned_index_),
value_type_,
+ ordered_);
}
protected:
@@ -570,6 +584,7 @@ class DictionaryBuilderBase : public ArrayBuilder {
BuilderType indices_builder_;
std::shared_ptr<DataType> value_type_;
bool ordered_ = false;
+ bool use_unsigned_index_ = false;
};
template <typename BuilderType>
@@ -581,10 +596,12 @@ class DictionaryBuilderBase<BuilderType, NullType> :
public ArrayBuilder {
start_int_size,
const std::shared_ptr<DataType>& value_type,
MemoryPool* pool = default_memory_pool(),
- int64_t alignment = kDefaultBufferAlignment, bool ordered = false)
+ int64_t alignment = kDefaultBufferAlignment, bool ordered = false,
+ bool use_unsigned_index = false)
: ArrayBuilder(pool, alignment),
indices_builder_(start_int_size, pool, alignment),
- ordered_(ordered) {}
+ ordered_(ordered),
+ use_unsigned_index_(use_unsigned_index) {}
explicit DictionaryBuilderBase(const std::shared_ptr<DataType>& value_type,
MemoryPool* pool = default_memory_pool(),
@@ -623,10 +640,11 @@ class DictionaryBuilderBase<BuilderType, NullType> :
public ArrayBuilder {
explicit DictionaryBuilderBase(const std::shared_ptr<Array>& dictionary,
MemoryPool* pool = default_memory_pool(),
int64_t alignment = kDefaultBufferAlignment,
- bool ordered = false)
+ bool ordered = false, bool use_unsigned_index
= false)
: ArrayBuilder(pool, alignment),
indices_builder_(pool, alignment),
- ordered_(ordered) {}
+ ordered_(ordered),
+ use_unsigned_index_(use_unsigned_index) {}
/// \brief Append a scalar null value
Status AppendNull() final {
@@ -678,7 +696,8 @@ class DictionaryBuilderBase<BuilderType, NullType> : public
ArrayBuilder {
Status FinishInternal(std::shared_ptr<ArrayData>* out) override {
ARROW_RETURN_NOT_OK(indices_builder_.FinishInternal(out));
- (*out)->type = dictionary((*out)->type, null(), ordered_);
+ (*out)->type = dictionary(MaybeUnsignedIndexType((*out)->type,
use_unsigned_index_),
+ null(), ordered_);
(*out)->dictionary = NullArray(0).data();
return Status::OK();
}
@@ -690,12 +709,15 @@ class DictionaryBuilderBase<BuilderType, NullType> :
public ArrayBuilder {
Status Finish(std::shared_ptr<DictionaryArray>* out) { return
FinishTyped(out); }
std::shared_ptr<DataType> type() const override {
- return ::arrow::dictionary(indices_builder_.type(), null(), ordered_);
+ return ::arrow::dictionary(
+ MaybeUnsignedIndexType(indices_builder_.type(), use_unsigned_index_),
null(),
+ ordered_);
}
protected:
BuilderType indices_builder_;
bool ordered_ = false;
+ bool use_unsigned_index_ = false;
};
} // namespace internal
diff --git a/cpp/src/arrow/builder.cc b/cpp/src/arrow/builder.cc
index 2190f1ce04..1d0a42099b 100644
--- a/cpp/src/arrow/builder.cc
+++ b/cpp/src/arrow/builder.cc
@@ -27,6 +27,7 @@
#include "arrow/util/checked_cast.h"
#include "arrow/util/hashing.h"
#include "arrow/util/logging_internal.h"
+#include "arrow/util/unreachable.h"
#include "arrow/visit_type_inline.h"
namespace arrow {
@@ -38,6 +39,40 @@ class MemoryPool;
using arrow::internal::checked_cast;
+namespace internal {
+
+/// Return the unsigned integer type of the same width when an unsigned
dictionary
+/// index type was requested (GH-37476).
+///
+/// The adaptive indices builder only ever produces signed integer types.
Dictionary
+/// indices are non-negative, so the signed and unsigned integer types of a
given width
+/// have identical memory layout and reporting one as the other is
value-preserving. The
+/// width stays adaptive, as it is for signed index types, and it widens on
the signed
+/// threshold: a uint8 index widens after 128 distinct values rather than the
256 a real
+/// uint8 could hold, so the extra bit does not delay widening.
+std::shared_ptr<DataType> MaybeUnsignedIndexType(
+ const std::shared_ptr<DataType>& index_type, bool use_unsigned_index) {
+ if (!use_unsigned_index) {
+ return index_type;
+ }
+ switch (index_type->id()) {
+ case Type::INT8:
+ return ::arrow::uint8();
+ case Type::INT16:
+ return ::arrow::uint16();
+ case Type::INT32:
+ return ::arrow::uint32();
+ case Type::INT64:
+ return ::arrow::uint64();
+ default:
+ // The adaptive index builder only ever produces signed int8/16/32/64,
so no
+ // other type reaches this point when an unsigned index was requested.
+ Unreachable("MaybeUnsignedIndexType: adaptive dictionary index type is
not signed");
+ }
+}
+
+} // namespace internal
+
// Generic int builder that delegates to the builder for a specific
// type. Used to reduce the number of template instantiations in the
// exact_index_type case below, to reduce build time and memory usage.
@@ -170,9 +205,10 @@ struct DictionaryBuilderCase {
using AdaptiveBuilderType = DictionaryBuilder<ValueType>;
using ExactBuilderType =
internal::DictionaryBuilderBase<TypeErasedIntBuilder, ValueType>;
+ const bool unsigned_index = is_unsigned_integer(index_type->id());
if (dictionary != nullptr) {
- out->reset(
- new AdaptiveBuilderType(dictionary, pool, kDefaultBufferAlignment,
ordered));
+ out->reset(new AdaptiveBuilderType(dictionary, pool,
kDefaultBufferAlignment,
+ ordered, unsigned_index));
} else if (exact_index_type) {
if (!is_integer(index_type->id())) {
return Status::TypeError("MakeBuilder: invalid index type ",
*index_type);
@@ -182,7 +218,8 @@ struct DictionaryBuilderCase {
} else {
auto start_int_size = index_type->byte_width();
out->reset(new AdaptiveBuilderType(start_int_size, value_type, pool,
- kDefaultBufferAlignment, ordered));
+ kDefaultBufferAlignment, ordered,
+ unsigned_index));
}
return Status::OK();
}
diff --git a/python/pyarrow/src/arrow/python/arrow_to_pandas.cc
b/python/pyarrow/src/arrow/python/arrow_to_pandas.cc
index 3cf6da31a8..a0300abbf9 100644
--- a/python/pyarrow/src/arrow/python/arrow_to_pandas.cc
+++ b/python/pyarrow/src/arrow/python/arrow_to_pandas.cc
@@ -1854,13 +1854,17 @@ class CategoricalWriter
}
Status WriteIndicesUniform(const ChunkedArray& data) {
- // For unsigned types, upcast to signed since pandas uses -1 for nulls
- // uint8 to int16, uint16 to int32, uint32 to int64, signed types unchanged
+ // For unsigned types, use a signed output since pandas uses -1 for nulls:
+ // uint8 to int16, uint16 to int32, uint32 to int64. uint64 also maps to
int64,
+ // which is safe because the indices are bounds-checked below the
dictionary length
+ // and so never reach the int64 range limit. Signed types are unchanged.
using OutputType = std::conditional_t<
std::is_same<T, uint8_t>::value, int16_t,
std::conditional_t<
std::is_same<T, uint16_t>::value, int32_t,
- std::conditional_t<std::is_same<T, uint32_t>::value, int64_t, T>>>;
+ std::conditional_t<
+ std::is_same<T, uint32_t>::value, int64_t,
+ std::conditional_t<std::is_same<T, uint64_t>::value, int64_t,
T>>>>;
const int npy_output_type = std::is_same<OutputType, int16_t>::value ?
NPY_INT16
: std::is_same<OutputType, int32_t>::value ?
NPY_INT32
: std::is_same<OutputType, int64_t>::value
@@ -2035,9 +2039,7 @@ Status MakeWriter(const PandasOptions& options,
PandasWriter::type writer_type,
CATEGORICAL_CASE(UInt8Type);
CATEGORICAL_CASE(UInt16Type);
CATEGORICAL_CASE(UInt32Type);
- case Type::UINT64:
- return Status::TypeError(
- "Converting UInt64 dictionary indices to pandas is not
supported.");
+ CATEGORICAL_CASE(UInt64Type);
default:
// Unreachable
ARROW_DCHECK(false);
diff --git a/python/pyarrow/tests/test_array.py
b/python/pyarrow/tests/test_array.py
index a26481c575..bc4e521dca 100644
--- a/python/pyarrow/tests/test_array.py
+++ b/python/pyarrow/tests/test_array.py
@@ -4583,3 +4583,42 @@ def test_dunders_checked_overflow():
arr ** pa.scalar(2, type=pa.int8())
with pytest.raises(pa.ArrowInvalid, match=error_match):
arr / (-arr)
+
+
[email protected]("index_type", [pa.int8(), pa.int16(), pa.int32(),
pa.int64(),
+ pa.uint8(), pa.uint16(), pa.uint32(),
+ pa.uint64()])
+def test_dictionary_array_preserves_index_type(index_type):
+ # GH-37476: an unsigned dictionary index type must be preserved, not
replaced by the
+ # signed integer type of the same width. Signed index types are kept as-is.
+ dict_type = pa.dictionary(index_type, pa.string())
+
+ arr = pa.array(["a", "b", None, "a"], type=dict_type)
+ assert arr.type == dict_type
+ assert arr.to_pylist() == ["a", "b", None, "a"]
+ arr.validate(full=True)
+
+ chunked = pa.chunked_array([["a", "b", "a"]], dict_type)
+ assert chunked.type == dict_type
+
+
[email protected]("start_type, widened_type", [(pa.int8(), pa.int16()),
+ (pa.uint8(),
pa.uint16())])
+def test_dictionary_array_index_width_adapts(start_type, widened_type):
+ # The index width adapts to the number of distinct values, as it does for
signed
+ # indices; only the signedness of the requested type is preserved.
+ values = [str(i) for i in range(200)]
+ arr = pa.array(values, type=pa.dictionary(start_type, pa.string()))
+ assert arr.type == pa.dictionary(widened_type, pa.string())
+ assert arr.to_pylist() == values
+
+
[email protected]
+def test_dictionary_uint64_index_to_pandas():
+ # GH-37476: uint64 dictionary indices are preserved, and converting them
to pandas
+ # maps the indices to int64 categorical codes, which is safe because the
indices are
+ # bounded by the dictionary length.
+ arr = pa.array(["a", "b", None, "a"], type=pa.dictionary(pa.uint64(),
pa.string()))
+ result = arr.to_pandas()
+ assert list(result.cat.categories) == ["a", "b"]
+ assert result.cat.codes.tolist() == [0, 1, -1, 0]
diff --git a/python/pyarrow/tests/test_pandas.py
b/python/pyarrow/tests/test_pandas.py
index 4b144448ce..0f35edd516 100644
--- a/python/pyarrow/tests/test_pandas.py
+++ b/python/pyarrow/tests/test_pandas.py
@@ -4149,21 +4149,14 @@ def test_dictionary_with_pandas():
d1 = pa.DictionaryArray.from_arrays(indices, dictionary)
d2 = pa.DictionaryArray.from_arrays(indices, dictionary, mask=mask)
- if index_type == 'uint64':
- # uint64 is not supported due to overflow risk (values > 2^63-1)
- with pytest.raises(TypeError,
- match="UInt64 dictionary indices"):
- d1.to_pandas()
- continue
-
pandas1 = d1.to_pandas()
# Pandas Categorical uses signed int codes. Arrow converts:
- # uint8 to int16, uint16 to int32, uint32 to int64, signed types
unchanged
+ # uint8 to int16, uint16 to int32, uint32 and uint64 to int64, signed
unchanged
if index_type == 'uint8':
compare_indices = indices.astype('int16')
elif index_type == 'uint16':
compare_indices = indices.astype('int32')
- elif index_type == 'uint32':
+ elif index_type in ('uint32', 'uint64'):
compare_indices = indices.astype('int64')
else:
compare_indices = indices
@@ -4179,7 +4172,7 @@ def test_dictionary_with_pandas():
signed_indices = indices.astype('int16')
elif index_type == 'uint16':
signed_indices = indices.astype('int32')
- elif index_type == 'uint32':
+ elif index_type in ('uint32', 'uint64'):
signed_indices = indices.astype('int64')
else:
signed_indices = indices