viirya commented on code in PR #56334:
URL: https://github.com/apache/spark/pull/56334#discussion_r3538352361


##########
sql/api/src/main/scala/org/apache/spark/sql/util/ArrowUtils.scala:
##########
@@ -38,6 +38,50 @@ private[sql] object ArrowUtils {
 
   // todo: support more types.
 
+  /**
+   * Check if a Spark DataType is supported by Arrow. This recursively checks 
complex types
+   * (Array, Struct, Map).
+   *
+   * Note: This checks compatibility with toArrowField(), not toArrowType(). 
Types like
+   * GeometryType, GeographyType, and VariantType are not supported by 
toArrowType() (which only
+   * handles primitive Arrow types), but ARE supported by toArrowField() which 
converts them to
+   * Arrow Struct representations with metadata. Since Arrow cache uses 
toArrowField() via
+   * toArrowSchema() to create the schema, these types are supported.
+   */
+  def isSupportedByArrow(dt: DataType): Boolean = {
+    dt match {
+      // Primitive types
+      case BooleanType | ByteType | ShortType | IntegerType | LongType | 
FloatType | DoubleType |
+          _: StringType | BinaryType | NullType =>
+        true
+
+      // Decimal
+      case _: DecimalType => true
+
+      // Temporal types
+      case DateType | TimestampType | TimestampNTZType | _: TimeType => true
+
+      // Interval types
+      case _: YearMonthIntervalType | _: DayTimeIntervalType | 
CalendarIntervalType => true

Review Comment:
   Follow-up filed: SPARK-58005 / https://github.com/apache/spark/pull/57088 
implements the plan above. It extends SPARK-57975's opt-in lossless encoding to 
`CalendarInterval` -- a struct of `(months, days, microseconds)`, the type's 
own layout, so the `* 1000` conversion and its overflow do not exist on that 
path -- and, independently, moves the overflow translation into 
`IntervalMonthDayNanoWriter` itself: the structured `DATETIME_OVERFLOW` is now 
raised by a catch scoped to the single `Math.multiplyExact` expression (the 
`TimestampNTZNanosWriter` pattern), so it structurally cannot re-label an 
unrelated `ArithmeticException` from upstream evaluation. Once both land, this 
PR will opt the cache into the lossless encoding and delete 
`withIntervalOverflowTranslation` entirely, which removes the mis-attribution 
you found along with the wrapper; your ANSI `DIVIDE_BY_ZERO` repro then 
surfaces `DIVIDE_BY_ZERO` regardless of interval columns in the schema.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to