malinjawi opened a new pull request, #13160:
URL: https://github.com/apache/gluten/pull/13160

   ### What changes were proposed in this pull request?
   
   Fix native partitioned Delta writes that exceed `maxRecordsPerFile` or roll 
files too early. Account rows per written chunk, fill the current file before 
opening another, and slice partition stripes that exceed its remaining 
capacity. Send each chunk to the statistics trackers for its destination file 
and release slices and remaining stripe batches on failure.
   
   Apply the fix to the Delta 3.3 and Delta 4.0 source paths. Add 13 shared 
backend regression cases and a Spark 3.5 `gluten-ut` case for null partition 
keys with file limits of 1 and 2. Declare its Delta runtime dependency in the 
test profile so standalone module runs exercise native writing. Include Delta 
suites in the x86 Spark 3.4–4.1 and enhanced Spark 4.0 CI selections.
   
   Related issue: #10215. Revives #12016; partition-column preservation remains 
separate in #12069.
   
   ### Why are the changes needed?
   
   The writer currently adds the original batch's row count to the last 
partition's file after writing every stripe. For example, three rows for 
partition A and one for B count as four rows against B's file, causing a later 
batch for B to roll the file early. A stripe larger than the file limit is also 
written without being split.
   
   ### Does this PR introduce any user-facing change?
   
   Native partitioned Delta writes honor `maxRecordsPerFile`, with file row 
counts and statistics updated for each written chunk. No public API or 
configuration change.
   
   ### How was this patch tested?
   
   The 13-case `DeltaNativeWriteLayoutSuite` covers partial files, partition 
boundaries across batches, oversized stripes, exact limits, empty batches, 
unlimited files, and file-open/write failures. Its integration matrix toggles 
statistics and optimized writes, asserts native execution, and checks physical 
Parquet rows, per-file statistics, duplicate-preserving readback, and history 
row counts using Spark as the read oracle.
   
   Completed on macOS arm64 with JDK 17:
   
   - Clean native build against pinned Velox 
`5cf370d2637d64feac516b49608daa9f95b0e120` passed.
   - Spark 4.1.1 / Delta 4.1.0 / Scala 2.13.17: the Delta package run passed 
301 ScalaTest tests and 15 JUnit tests. The final shared layout suite also 
passed 13/13 after correcting its handling of URI-encoded filenames.
   - Regression check: the old writer implementation produced six expected 
failures; restoring the fix passed all 13 layout cases.
   - Spark 3.5.5 / Delta 3.3.2 / Scala 2.12.18: clean reactor build passed. All 
277 existing Delta cases passed in the package run; after correcting the test 
path helper, the final layout suite passed 13/13 in a focused rerun. The 
separate `gluten-ut` null-partition test passed 1/1.
   
   Both package runs used the normal CI tag exclusions and reported 262 ignored 
cases. The Spark 3.5 package initially failed the eight layout integration 
cases because the new helper treated URI-encoded filenames as literal paths; 
the focused rerun validates that correction. No existing Delta suite failed in 
the fresh test warehouse.
   
   Scala/POM formatting, license-header checks, workflow selection and shell 
syntax checks, and `git diff --check` passed.
   
   ### Was this patch authored or co-authored using generative AI tooling?
   
   Generated-by: OpenAI Codex (GPT-6)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to