felipepessoto opened a new pull request, #13136:
URL: https://github.com/apache/gluten/pull/13136

   <!--
   Thank you for submitting a pull request! Here are some tips:
   
   1. For first-time contributors, please read our contributing guide:
      https://github.com/apache/gluten/blob/main/CONTRIBUTING.md
   2. If necessary, create a GitHub issue for discussion beforehand to avoid 
duplicate work.
   3. If the PR is specific to a single backend, include [VL] or [CH] in the PR 
title to indicate the
      Velox or ClickHouse backend, respectively.
   4. If the PR is not ready for review, please mark it as a draft.
   -->
   
   ## What changes are proposed in this pull request?
   
   <!--
   Provide a clear and concise description of the changes introduced in this PR.
   Ensure the PR description aligns with the code changes, especially after 
updates.
   If applicable, include "Fixes #<GitHub_Issue_ID>" to automatically close the 
corresponding issue
   when the PR is merged.
   -->
   
   Separate the physical reader schema from the logical scan-output schema in 
the Velox plan converter. A metadata-only query such as `SELECT DISTINCT 
_metadata.file_name` over a Parquet table containing a user `file_name BIGINT` 
must not request that physical field as `VARCHAR` before the connector installs 
file metadata constants.
   
   - Without an explicit split table schema, build 
`HiveTableHandle::dataColumns()` from **regular scan slots by role and index**. 
Partition, synthesized, and generated row-index outputs remain available to the 
scan but do not become requested physical fields. Metadata-only projections 
retain a non-null empty physical schema.
   - With an explicit table schema, retain genuine physical names, types, and 
ordinals, applying only the existing partition exclusion. Do not remove a real 
field just because a synthetic output has the same name.
   - Preserve logical output slots, assignment roles, filter binding, split 
metadata, and the existing Spark fallback for simultaneous real and same-named 
metadata outputs. No production Scala/Substrait protocol changes or permissive 
cast changes.
   - Add programmatic native converter/Parquet regressions and a bounded Spark 
4.1 regression matrix for all six file constants, incompatible same-name 
physical types, aliases/whole structs, filters, partitions, case normalization, 
and explicit positional schemas. The native fixture now forwards parsed split 
metadata and writes actual Parquet fixtures.
   
   This is independent of the generated Delta DV metadata work in #13123; it 
does not change that feature's protocol or implement the separate Parquet 
date-widening fix.
   
   ## How was this patch tested?
   
   <!--
   Describe how the changes were tested, if applicable.
   Include new tests to validate the functionality, if necessary.
   For UI-related changes, attach screenshots to demonstrate the updates.
   -->
   
   **Draft: compilation and runtime validation are pending.**
   
   Completed:
   - Repository Scala formatter (`dev/format-scala-code.sh`) through the Maven 
wrapper in the Gluten JDK 17 container.
   - clang-format **15.0.7** on the changed C++ files, including a final 
dry-run check.
   - Repository CI license-header checker on all changed files, and `git diff 
--check`. The documented `dev/check.py header` wrapper refers to a missing 
`dev/license-header.py`, so the existing 
`.github/workflows/util/license-header.py` checker was used directly.
   
   Added but **not yet compiled/executed**:
   - Six native converter/Parquet tests: regular-slot selection, empty/untagged 
schemas, explicit physical schema preservation, per-file metadata and row 
counts (including an empty file), metadata/regular filters, genuine 
regular-type rejection, and positional mapping.
   - Twelve Spark 4.1 metadata regression cases, checking actual metadata 
values and native scans for supported projections, while explicitly retaining 
same-name combined-output fallback.
   
   An exact-source native build was attempted in an isolated 
`apache/gluten:centos-9-jdk17` container. CMake stopped before compilation 
because the IBM-pinned Velox dependency source/build tree is not prepared. No 
sibling PR native binaries were substituted. The native and JVM regression 
suites, baseline/fixed red-green reproduction, and the original twelve Delta 
aliasing failures have not been run locally. Native PR CI and focused runtime 
regressions remain acceptance gates before marking this ready.
   
   ## Was this patch authored or co-authored using generative AI tooling?
   
   <!--
   If generative AI tooling has been used in the process of authoring this 
patch, please include the
   phrase: 'Generated-by: ' followed by the name of the tool and its version.
   If no, write 'No'.
   Please refer to the [ASF Generative Tooling 
Guidance](https://www.apache.org/legal/generative-tooling.html) for details.
   -->
   
   Generated-by: GitHub Copilot 1.0.87-0 (GPT-6 Astra)


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to