felipepessoto opened a new pull request, #13123:
URL: https://github.com/apache/gluten/pull/13123

   <!--
   Thank you for submitting a pull request! Here are some tips:
   
   1. For first-time contributors, please read our contributing guide:
      https://github.com/apache/gluten/blob/main/CONTRIBUTING.md
   2. If necessary, create a GitHub issue for discussion beforehand to avoid 
duplicate work.
   3. If the PR is specific to a single backend, include [VL] or [CH] in the PR 
title to indicate the
      Velox or ClickHouse backend, respectively.
   4. If the PR is not ready for review, please mark it as a draft.
   -->
   
   ## What changes are proposed in this pull request?
   
   <!--
   Provide a clear and concise description of the changes introduced in this PR.
   Ensure the PR description aligns with the code changes, especially after 
updates.
   If applicable, include "Fixes #<GitHub_Issue_ID>" to automatically close the 
corresponding issue
   when the PR is merged.
   -->
   
   Partially addresses #12743. This is a **standalone draft targeting main**, 
exploring a native alternative/follow-up to the JVM correctness fallback in 
#13042. It does not resolve the full issue, is not stacked on #13042, and does 
not modify that PR.
   
   **Native C++ compilation and the new end-to-end DV tests remain unverified 
locally.** This draft needs Linux native/Delta CI before it can be considered 
ready; it makes no performance claim.
   
   For `spark.databricks.delta.deletionVectors.useMetadataRowIndex=false`:
   
   - Introduce distinct Substrait column kinds for generated Delta row indexes 
and deleted-row flags. Recognize exact Delta names without requiring 
generated-metadata markers, checking both output and required schema.
   - Generate absolute file positions through Velox's row-index reader. A 
private BIGINT channel supplies positions for bitmap lookup; the public 
deleted-status output remains TINYINT. No extra native plan operator is 
inserted, preserving the existing metrics mapping.
   - **Mark, do not prematurely remove, rows** when a deleted-status output is 
requested. `IF_CONTAINED` marks bitmap members; `IF_NOT_CONTAINED` marks 
non-members. An absent DV yields `0`, while a present empty inverse bitmap 
yields `1` for every row.
   - Preserve consumer predicates and projections, and keep generated-field 
predicates out of physical Parquet pushdown. Enforce generated-key dynamic 
filters after materialization, including Velox's join-replacement optimization, 
empty batches and preloaded-source handoff. Restrict legacy predicate/column 
stripping to Delta's injected unary shape so an ordinary scan in another join 
branch cannot strip generated outputs.
   - Remove Spark 4.0/4.1's generic classification of the deleted-status name 
as a row-index column. Delta-specific classification stays in the Delta 
transformer.
   
   The existing metadata-row-index=true native masking path, executor-deferred 
checksum-validating DV payload reads, memoization, metrics, and authoritative 
PreparedDeltaFileIndex AddFiles are retained. The specialized Tahoe metadata 
route still supplies inverse filter semantics. This builds on the handoff 
merged in #12836; the closed, unmerged native range-read proposal #12867 is not 
included.
   
   JVM fallback remains for Spark 3.4 DV scans, CDF scans touching DVs, the DML 
configuration escape hatch, DV-bearing scans without a deleted-status output, 
and unsupported generated-field/schema shapes (including required-only fields 
absent from output, ambiguous/malformed fields, bucketed scans and 
mapped/partition-name collisions). DV-free row-index-only scans are supported.
   
   Adds shared Delta 3.3/4.0 integration cases with native-plan assertions, 
native connector/conversion regressions, a gluten-ut contract/shim suite, and 
related Delta documentation. Coverage includes the nine DV-free field/reader 
combinations, full value/index association, combined fields, raw marked rows, 
inverse/mixed-file semantics, row-group pruning, split offsets, small batches, 
name mapping, mixed-scan projections, repeated DML and fallback. Native 
unique-key inner/semi-join cases cover generated keys and split preloading; 
inner cases assert `replacedWithDynamicFilterRows > 0` so correctness cannot be 
supplied only by a remaining hash lookup.
   
   **The known-failures baseline is deliberately unchanged until CI provides 
evidence.** Rebuild and deploy matching native libraries and JVM artifacts for 
the new column-kind protocol.
   
   ## How was this patch tested?
   
   <!--
   Describe how the changes were tested, if applicable.
   Include new tests to validate the functionality, if necessary.
   For UI-related changes, attach screenshots to demonstrate the updates.
   -->
   
   Local environment: Windows, Git Bash, Temurin JDK 17; WSL Ubuntu 24.04 was 
also used for available checks.
   
   **Passed:**
   
   - Maven-wrapper Spark 3.5 / Scala 2.12 / Delta 3.3 production and 
test-source compilation. The final compilation was `bash .\build\mvn -B -q -pl 
gluten-ut/test -am test-compile 
-Pdelta,spark-ut,backends-velox,spark-3.5,scala-2.12,fast-build -DskipTests`. A 
matching reactor `install -DskipTests` also completed during validation.
   - `org.apache.gluten.sql.shims.DeltaMetadataColumnSuite`: **2 tests 
passed**, using ScalaTest with a Windows-safe classpath JAR generated from this 
worktree and Maven-resolved dependencies. The repository build-info script 
supplied the resource whose Maven generation is Unix-only.
   - Constructor-level registration check: **all 27 Delta 3.3 Handoff tests 
registered**, including the mapping, DML, fallback and mixed-scan cases. This 
is a registration check, not a runtime DV test pass.
   - `dev/format-scala-code.sh` and scoped Spotless check across all changed 
Scala files/profiles.
   - `dev/format-cpp-code.sh` in an isolated Linux validation copy, plus 
explicit clang-format **15.0.7** formatting and `--dry-run --Werror` checks for 
every touched C++ file, including `.cpp` files not selected by the script.
   - Actual CI license-header fixer/checker, 
`.github/workflows/util/license-header.py`, on every changed file; `git diff 
--check`. The documented `dev/check.py header main --fix` entry point was 
attempted but refers to a missing `dev/license-header.py`, so the existing CI 
implementation was used instead.
   
   **Blocked / not claimed as passing:**
   
   - Native build attempt, after normalizing shell line endings only in the 
isolated validation copy: `dev/builddeps-veloxbe.sh build_gluten_cpp 
--enable_s3=OFF --enable_gcs=OFF --enable_hdfs=OFF --enable_abfs=OFF 
--build_tests=ON` reached CMake and stopped at `cmake: command not found`. The 
WSL native compiler/Ninja/dependency toolchain is not provisioned. **No C++ 
compile or native test pass is claimed.**
   - Delta Handoff suite runtime startup aborts at missing 
`linux/amd64/gluten.dll`; **zero integration tests executed**.
   - Spark 4.0 / Scala 2.13 / JDK17 wrapper attempt with 
`-Dmaven.compiler.release=17` reaches the backend module and fails on six 
existing `src-delta40` symlink placeholders parsed as Scala source. Full Delta 
4.0 compilation/runtime is not claimed. A narrowed standalone test-source 
attempt also lacked the vendored Delta test utilities normally compiled by that 
module.
   - An initial reactor `test` attempt additionally encountered the existing 
Windows symlink-loop assertion in `JniLibLoaderTest`; the focused ScalaTest 
execution above avoids running unrelated reactor tests.
   
   Linux follow-up targets are `velox_delta_read_test` and 
`velox_plan_conversion_test`, plus `DeltaDeletionVectorHandoffSuite` and the 
existing Delta file-format/DV/DML/CDF regressions under both Spark 3.5/Delta 
3.3 and Spark 4.0/Delta 4.0. The native join-replacement and preloading 
assertions are included but have not been executed locally.
   
   ## Was this patch authored or co-authored using generative AI tooling?
   
   <!--
   If generative AI tooling has been used in the process of authoring this 
patch, please include the
   phrase: 'Generated-by: ' followed by the name of the tool and its version.
   If no, write 'No'.
   Please refer to the [ASF Generative Tooling 
Guidance](https://www.apache.org/legal/generative-tooling.html) for details.
   -->
   
   Generated-by: GitHub Copilot CLI 1.0.87-0 (GPT-6 Astra, gpt-6-astra)


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to