andygrove opened a new pull request, #6740:
URL: https://github.com/apache/datafusion-comet/pull/6740

   ## Which issue does this PR close?
   
   Closes #6739.
   
   ## Rationale for this change
   
   Spark 4.2.0 raises `long overflow` when SECOND or MILLISECOND truncation 
falls below the smallest timestamp, like every other unit (SPARK-56663). Spark 
4.1 and earlier wrap the Long subtraction instead. The fixture added in #5956 
expects the wrap, so `trunc_timestamp.sql` fails on the Spark 4.2 profile. 
Comet's native kernel also wraps on every version, so on Spark 4.2 Comet 
returns a timestamp in the year 294247 where Spark raises.
   
   The pull request tier only runs Spark 4.1, so the failure first showed up on 
#6664, which runs every profile. It also turns the 4.2 nightly red.
   
   ## What changes are included in this PR?
   
   - `TruncTimestamp` gets a `wrap_second_millisecond_overflow` proto field. 
The serde sets it to `!isSpark42Plus`, the same way `Mode.normalize_neg_zero` 
tracks the running Spark version. The proto default, `false`, is the 4.2 
behavior.
   - When the flag is off, the native kernel uses checked subtraction for 
SECOND and MILLISECOND, as it already does for every other unit. Dictionary 
input then keeps key masking for those two units when a value is near the lower 
bound, so an unused dictionary entry at `Long.MinValue` still never raises.
   - The SECOND/MILLISECOND query moves out of `trunc_timestamp.sql` into a 
pair of version-gated fixtures. `trunc_timestamp_fine_overflow.sql` covers 
Spark 4.1 and earlier, where both Spark and Comet wrap. 
`trunc_timestamp_fine_overflow_spark42.sql` covers 4.2, where both raise. It 
also runs a native `date_trunc` query over valid input, so the errors cannot 
come from a fallback to Spark.
   - A note in the `date_trunc` expression audit.
   
   ## How are these changes tested?
   
   - New Rust test `test_timestamp_trunc_fine_spark42_overflow` covers the 
error, the first representable boundary and NULLs for both units across 
timezones. The dictionary unused-extremes test and the poisoned-NULL test now 
run in both modes.
   - `CometSqlFileTestSuite trunc_timestamp` and `CometTemporalExpressionSuite 
date_trunc` pass locally on Spark 4.2, 4.1 and 3.5, 27 tests each.
   - Mutation check: with the serde forced to always wrap, 
`trunc_timestamp_fine_overflow_spark42.sql` fails on Spark 4.2 with "Expected 
Comet to throw an error matching 'long overflow' but query succeeded".
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to