adriangb commented on PR #11038: URL: https://github.com/apache/arrow-rs/pull/11038#issuecomment-5872798297
@Jefffrey thanks for the review. I applied both suggestions and added the note about `None` in 63df2e0329. Regarding Spark: indeed it does what Java does. I ran Spark 4.2.0 locally on the same cases as the table in the PR description. `CAST(TIMESTAMP_NTZ ... AS TIMESTAMP)`, `to_utc_timestamp`, `convert_timezone` and a string cast all give the same result, with ANSI mode on and off. All times are UTC: | Case | Spark 4.2.0 | PostgreSQL 17 / DuckDB / this PR | | --- | --- | --- | | `America/New_York` `2024-11-03 01:30` — **ambiguous** | `05:30` (earlier) | `06:30` (later) | | `America/Havana` `2024-11-03 00:00` — **ambiguous** | `04:00` (earlier) | `05:00` (later) | | `Australia/Sydney` `2024-04-07 02:30` — **ambiguous** | `2024-04-06 15:30` (earlier) | `2024-04-06 16:30` (later) | | `America/New_York` `2024-03-10 02:30` — gap | `07:30` | `07:30` | | `America/Sao_Paulo` `2018-11-04 00:00` — gap | `03:00` | `03:00` | | `Australia/Sydney` `2024-10-06 02:30` — gap | `2024-10-05 16:30` | `2024-10-05 16:30` | | `Australia/Lord_Howe` `2024-10-06 02:15` — gap | `2024-10-05 15:45` | `2024-10-05 15:45` | | `Pacific/Chatham` `2024-09-29 03:00` — gap | `2024-09-28 14:15` | `2024-09-28 14:15` | So Spark agrees on every gap and disagrees on every ambiguous reading. The source confirms it: the cast calls [`convertTz`](https://github.com/apache/spark/blob/b1b2685058d6359e00842cbd4c723f4ce3810709/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/Cast.scala#L901), which calls [`LocalDateTime.atZone`](https://github.com/apache/spark/blob/b1b2685058d6359e00842cbd4c723f4ce3810709/sql/api/src/main/scala/org/apache/spark/sql/catalyst/util/SparkDateTimeUtils.scala#L198). That is `ZonedDateTime.of`, which takes the earlier offset in an overlap and shifts forward in a gap. Spark never raises an error or returns NULL for these readings. So the SQL engines split on the ambiguous case: PostgreSQL and DuckDB take the later instant, Spark takes the earlier one. But, as far as I was able to tell, no Spark-compatible DataFusion implementer reaches this code. Comet handles the naive-to-zoned cast in its own code ([cast.rs](https://github.com/apache/datafusion-comet/blob/764936187/native/spark-expr/src/conversion_funcs/cast.rs#L395)), which takes the earlier instant. That code runs before Comet falls back to arrow's cast. In datafusion-spark, to_utc_timestamp, from_utc_timestamp and spark_cast do not call it either. And compared with Spark, readings in a gap go from an error or NULL to Spark's answer. Ambiguous readings did not give Spark's answer before this PR either. So this PR does not make anything worse for Spark users, because no Spark-compatible engine reaches this kernel for this cast. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
