luis4a0 opened a new pull request, #13154: URL: https://github.com/apache/gluten/pull/13154
## What changes are proposed in this pull request? Offload Spark 4.1 PMOD for BYTE, SHORT, INT, LONG, FLOAT, and DOUBLE through Velox's captured-mode `pmod_with_mode` special form. DECIMAL remains on Spark. - Preserve the mode captured by the Catalyst expression, including execution after session settings change. - Serialize both native Boolean arguments in the Substrait function signature while retaining Spark's result type and nullability. - Preserve divisor-first NULL, zero-divisor, and composed-error behavior instead of using eager scalar argument evaluation. - Check the exact loaded native capability even when general native validation is disabled. An older dependency retains Spark execution rather than using the incompatible legacy signature. - Keep nonnullable ANSI compositions with potentially throwing computed dividends on Spark: Spark's interpreter and generated code choose different error precedence for these expressions. - Preserve other backends through the existing generic transformation behavior behind a backend-specific hook. ### Native dependency prerequisite This requires the companion Velox change providing `pmod_with_mode`: https://github.com/facebookincubator/velox/pull/19231 Gluten's currently pinned public Velox tag, `dft-2026_09_25`, does not contain it. Local integration was tested with that public dependency plus the companion change. The stock dependency was also exercised to verify safe Spark fallback, including with general native validation disabled. The positive native-offload tests require a compatible dependency update before this PR can merge; this PR does not add a temporary dependency patch or weaken those assertions. ## How was this patch tested? - Six native registration/regression tests, including PMOD captured-mode changes and dividend suppression. - Six transformation/validation tests covering every primitive type, complete native argument signatures, captured flags, mandatory capability checking, and unsupported cases. - Twelve Spark 4.1 integration tests pass on JDK 17 and JDK 21. These compare results with Spark and require actual native PMOD execution for supported cases, including LEGACY short circuit with two nonnullable operands. - Integration cases cover operand signs, each integer width, wrapping boundaries, floating special values and signed zero, zero divisors, NULLs, mixed numeric coercion, changes after analysis and physical planning, and explicit Spark fallback. - Full physical Spark execution verifies the distinct `REMAINDER_BY_ZERO` and `USER_RAISED_EXCEPTION` outcomes under generated and interpreted evaluation for the ambiguous composition. ## Was this patch authored or co-authored using generative AI tooling? Yes. Co-authored by Luis Penaranda and GitHub Copilot CLI. Generated-by: GitHub Copilot CLI 1.0.87-0 (GPT-6 Astra) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
