andygrove opened a new issue, #6330:
URL: https://github.com/apache/datafusion-comet/issues/6330

   ### Describe the bug
   
   Since #4761, `TimestampTruncExpr` stamps its output with the session 
timezone and declares `Timestamp(Microsecond, <session timezone>)` as its type 
(`native/spark-expr/src/datetime_funcs/timestamp_trunc.rs:93`, 
`native/spark-expr/src/kernels/temporal.rs:132` and `:834`). Everything else in 
a native plan represents `TimestampType` as `Timestamp(Microsecond, "UTC")`. 
Arrow's comparison kernels require identical types, so comparing a truncated 
timestamp with any other timestamp fails.
   
   On its own, that would only matter with 
`spark.comet.expression.TruncTimestamp.allowIncompatible=true`. But 
`CometTruncTimestamp` treats `Etc/UTC` as UTC 
(`spark/src/main/scala/org/apache/comet/serde/datetime.scala:610`) and runs the 
native path by default there, and the output is then labelled `Etc/UTC`. 
`Etc/UTC` is the JVM's default timezone on Linux hosts and containers whose 
`/etc/localtime` points at it, which is the default on Ubuntu and Debian 
images. That makes it Spark's default `spark.sql.session.timeZone` there too.
   
   In an `Etc/UTC` session, comparing `date_trunc(...)` with another timestamp 
fails. That covers `=`, `>=`, `<=`, `BETWEEN`, `<=>`, `nullif` and join 
conditions. `CASE` and `coalesce` panic instead (#6327). `IF`, `IN` lists, 
`greatest`, equi-join keys, `UNION`, aggregates, sorts and windows all work.
   
   ### Steps to reproduce
   
   On `main` at `764936187`, with the default config, on Spark 3.5 and 4.1:
   
   ```sql
   CREATE TABLE events USING parquet AS
   SELECT * FROM VALUES (TIMESTAMP'2024-01-15T18:30:45Z'), 
(TIMESTAMP'2024-06-30T23:30:00Z') AS v(ts);
   
   SET spark.sql.session.timeZone=Etc/UTC;
   SELECT ts FROM events WHERE date_trunc('DAY', ts) >= TIMESTAMP'2024-06-01 
00:00:00';
   ```
   
   Spark returns `2024-06-30 23:30:00`. Comet fails with `Invalid argument 
error: Invalid comparison operation: Timestamp(µs, "Etc/UTC") >= Timestamp(µs, 
"UTC")`. The same query works with `spark.sql.session.timeZone=UTC`. With 
`allowIncompatible=true` it fails the same way in every other zone, for example 
`Timestamp(µs, "Asia/Tokyo") >= Timestamp(µs, "UTC")`.
   
   ### Expected behavior
   
   The same result as Spark.
   
   ### Additional context
   
   This has been in place since 1.0.0. #5556 already had to accept `Etc/UTC` 
and `UTC` as equivalent in the Python runner because of this label. Rather than 
teaching each consumer, could `TimestampTruncExpr` do the truncation in the 
session timezone but keep the input's label on the output, and return the 
child's type from `data_type()`? Declared and actual types would still agree, 
so the `RowConverter` mismatch #4761 fixed stays fixed. #5956 changes the same 
kernel. Normalizing the session timezone (#6329) would hide this for `Etc/UTC`, 
but the label would still leak in every other zone under `allowIncompatible`.
   
   The incompatibility reason for non-UTC `date_trunc` still cites #2649, which 
#4761 closed. The remaining non-UTC gaps are #5633 and the chrono-tz horizon 
described in the datetime compatibility guide.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to