viirya commented on code in PR #56334:
URL: https://github.com/apache/spark/pull/56334#discussion_r3554429297
##########
sql/api/src/main/scala/org/apache/spark/sql/util/ArrowUtils.scala:
##########
@@ -38,6 +38,50 @@ private[sql] object ArrowUtils {
// todo: support more types.
+ /**
+ * Check if a Spark DataType is supported by Arrow. This recursively checks
complex types
+ * (Array, Struct, Map).
+ *
+ * Note: This checks compatibility with toArrowField(), not toArrowType().
Types like
+ * GeometryType, GeographyType, and VariantType are not supported by
toArrowType() (which only
+ * handles primitive Arrow types), but ARE supported by toArrowField() which
converts them to
+ * Arrow Struct representations with metadata. Since Arrow cache uses
toArrowField() via
+ * toArrowSchema() to create the schema, these types are supported.
+ */
+ def isSupportedByArrow(dt: DataType): Boolean = {
+ dt match {
+ // Primitive types
+ case BooleanType | ByteType | ShortType | IntegerType | LongType |
FloatType | DoubleType |
+ _: StringType | BinaryType | NullType =>
+ true
+
+ // Decimal
+ case _: DecimalType => true
+
+ // Temporal types
+ case DateType | TimestampType | TimestampNTZType | _: TimeType => true
+
+ // Interval types
+ case _: YearMonthIntervalType | _: DayTimeIntervalType |
CalendarIntervalType => true
Review Comment:
Done. With SPARK-57975 and SPARK-58005 merged, this PR now opts the cache
into the lossless encodings (`losslessInternalTypes = true` at the
schema-construction sites) and deletes `withIntervalOverflowTranslation`
entirely -- on the cache path there is no longer a nanosecond conversion that
can overflow, so the mis-attribution you found is gone along with the wrapper:
your ANSI `DIVIDE_BY_ZERO` repro surfaces `DIVIDE_BY_ZERO` regardless of
interval columns. The overflow diagnostic test is replaced by a full-domain
round-trip test (`microseconds = Long.MaxValue` etc., top-level and nested),
and the docs now state the cache matches the default serializer's full value
domain for this type.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]