andygrove opened a new pull request, #6312:
URL: https://github.com/apache/datafusion-comet/pull/6312

   ## Which issue does this PR close?
   
   Closes #6301.
   
   ## Rationale for this change
   
   The nightly failed on Spark 3.4 in `ParquetReadV1Suite` "a file without ids 
next to a file with ids is checked on its own", which #6116 added. The 
assertion that failed was on the Spark reference error, not Comet's, and the 
same test passed on Spark 3.4 in the previous nightly.
   
   The test writes each side with `sparkContext.parallelize` on the five-core 
test session, so each write leaves an empty `part-00000` next to two one-row 
files. The read packs the six files into three tasks: task 0 `[ids, ids]`, task 
1 `[no ids, no ids]` and task 2 `[empty with ids, empty without ids]`. On Spark 
3.x, tasks 1 and 2 fail with different exceptions:
   
   - Task 1 fails on its first file and raises the bare `RuntimeException` from 
`ParquetReadSupport`.
   - Task 2 reads the file without ids after an empty file. Its error is raised 
inside the `try { hasNext }` in `FileScanRDD.nextIterator`, which wraps it in 
`SparkException("Encountered error while reading file ...")`.
   
   On 3.x the `DAGScheduler` wraps whichever task fails first in "Job aborted", 
and `isMissingFieldIdsError` looks only at `getCause`. The test passes when 
task 1 fails first and fails when task 2 does. The nightly log reports `Task 2 
in stage 2468.0 failed`. Spark 3.5 has the same exposure. Spark 4.x wraps every 
read error once in `FAILED_READ_FILE` and does not add "Job aborted", so the 
test always passes there. Comet raises the same exception from every task on 
every version.
   
   ## What changes are included in this PR?
   
   The test writes each side with `.repartition(1)`, which is how the suite 
already writes a single file. With one file per side there is no empty file, 
only one task fails, and Spark raises the bare `RuntimeException` every time. 
The assertion keeps Spark's own `getCause` form from `ParquetFieldIdIOSuite`. 
The test comment explains why each side is one file.
   
   `branch-1.1` has the same test through #6266 and will need the same change.
   
   ## How are these changes tested?
   
   The change is to the test itself. The race cannot be forced from the suite, 
so the fix was checked in two ways:
   
   - A plain Spark program (no Comet) repeats the test's writes and drains each 
read partition on its own. On Spark 3.4.3 and 3.5.9 it shows the three-task 
layout and the two exception shapes above. With `.repartition(1)` it shows two 
single-file tasks, and the one that fails raises the bare `RuntimeException` in 
30 of 30 runs on each version.
   - The `ParquetReadV1Suite` tests with "ids" in their name (7 tests, 
including this one) pass locally on the default profile (Spark 4.1) and with 
`-Pspark-3.4`. On Spark 3.4 the test now writes two files, of 454 and 476 
bytes: the same two-row files as the plain Spark program, and no empty file.
   
   The `run-all-spark-profiles` label is on this PR, so the Comet suites run 
against Spark 3.4, 3.5, 4.0 and 4.2 before it lands.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to