andygrove opened a new issue, #6330:
URL: https://github.com/apache/datafusion-comet/issues/6330
### Describe the bug
Since #4761, `TimestampTruncExpr` stamps its output with the session
timezone and declares `Timestamp(Microsecond, <session timezone>)` as its type
(`native/spark-expr/src/datetime_funcs/timestamp_trunc.rs:93`,
`native/spark-expr/src/kernels/temporal.rs:132` and `:834`). Everything else in
a native plan represents `TimestampType` as `Timestamp(Microsecond, "UTC")`.
Arrow's comparison kernels require identical types, so comparing a truncated
timestamp with any other timestamp fails.
On its own, that would only matter with
`spark.comet.expression.TruncTimestamp.allowIncompatible=true`. But
`CometTruncTimestamp` treats `Etc/UTC` as UTC
(`spark/src/main/scala/org/apache/comet/serde/datetime.scala:610`) and runs the
native path by default there, and the output is then labelled `Etc/UTC`.
`Etc/UTC` is the JVM's default timezone on Linux hosts and containers whose
`/etc/localtime` points at it, which is the default on Ubuntu and Debian
images. That makes it Spark's default `spark.sql.session.timeZone` there too.
In an `Etc/UTC` session, comparing `date_trunc(...)` with another timestamp
fails. That covers `=`, `>=`, `<=`, `BETWEEN`, `<=>`, `nullif` and join
conditions. `CASE` and `coalesce` panic instead (#6327). `IF`, `IN` lists,
`greatest`, equi-join keys, `UNION`, aggregates, sorts and windows all work.
### Steps to reproduce
On `main` at `764936187`, with the default config, on Spark 3.5 and 4.1:
```sql
CREATE TABLE events USING parquet AS
SELECT * FROM VALUES (TIMESTAMP'2024-01-15T18:30:45Z'),
(TIMESTAMP'2024-06-30T23:30:00Z') AS v(ts);
SET spark.sql.session.timeZone=Etc/UTC;
SELECT ts FROM events WHERE date_trunc('DAY', ts) >= TIMESTAMP'2024-06-01
00:00:00';
```
Spark returns `2024-06-30 23:30:00`. Comet fails with `Invalid argument
error: Invalid comparison operation: Timestamp(µs, "Etc/UTC") >= Timestamp(µs,
"UTC")`. The same query works with `spark.sql.session.timeZone=UTC`. With
`allowIncompatible=true` it fails the same way in every other zone, for example
`Timestamp(µs, "Asia/Tokyo") >= Timestamp(µs, "UTC")`.
### Expected behavior
The same result as Spark.
### Additional context
This has been in place since 1.0.0. #5556 already had to accept `Etc/UTC`
and `UTC` as equivalent in the Python runner because of this label. Rather than
teaching each consumer, could `TimestampTruncExpr` do the truncation in the
session timezone but keep the input's label on the output, and return the
child's type from `data_type()`? Declared and actual types would still agree,
so the `RowConverter` mismatch #4761 fixed stays fixed. #5956 changes the same
kernel. Normalizing the session timezone (#6329) would hide this for `Etc/UTC`,
but the label would still leak in every other zone under `allowIncompatible`.
The incompatibility reason for non-UTC `date_trunc` still cites #2649, which
#4761 closed. The remaining non-UTC gaps are #5633 and the chrono-tz horizon
described in the datetime compatibility guide.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]