andygrove opened a new pull request, #6312:
URL: https://github.com/apache/datafusion-comet/pull/6312
## Which issue does this PR close?
Closes #6301.
## Rationale for this change
The nightly failed on Spark 3.4 in `ParquetReadV1Suite` "a file without ids
next to a file with ids is checked on its own", which #6116 added. The
assertion that failed was on the Spark reference error, not Comet's, and the
same test passed on Spark 3.4 in the previous nightly.
The test writes each side with `sparkContext.parallelize` on the five-core
test session, so each write leaves an empty `part-00000` next to two one-row
files. The read packs the six files into three tasks: task 0 `[ids, ids]`, task
1 `[no ids, no ids]` and task 2 `[empty with ids, empty without ids]`. On Spark
3.x, tasks 1 and 2 fail with different exceptions:
- Task 1 fails on its first file and raises the bare `RuntimeException` from
`ParquetReadSupport`.
- Task 2 reads the file without ids after an empty file. Its error is raised
inside the `try { hasNext }` in `FileScanRDD.nextIterator`, which wraps it in
`SparkException("Encountered error while reading file ...")`.
On 3.x the `DAGScheduler` wraps whichever task fails first in "Job aborted",
and `isMissingFieldIdsError` looks only at `getCause`. The test passes when
task 1 fails first and fails when task 2 does. The nightly log reports `Task 2
in stage 2468.0 failed`. Spark 3.5 has the same exposure. Spark 4.x wraps every
read error once in `FAILED_READ_FILE` and does not add "Job aborted", so the
test always passes there. Comet raises the same exception from every task on
every version.
## What changes are included in this PR?
The test writes each side with `.repartition(1)`, which is how the suite
already writes a single file. With one file per side there is no empty file,
only one task fails, and Spark raises the bare `RuntimeException` every time.
The assertion keeps Spark's own `getCause` form from `ParquetFieldIdIOSuite`.
The test comment explains why each side is one file.
`branch-1.1` has the same test through #6266 and will need the same change.
## How are these changes tested?
The change is to the test itself. The race cannot be forced from the suite,
so the fix was checked in two ways:
- A plain Spark program (no Comet) repeats the test's writes and drains each
read partition on its own. On Spark 3.4.3 and 3.5.9 it shows the three-task
layout and the two exception shapes above. With `.repartition(1)` it shows two
single-file tasks, and the one that fails raises the bare `RuntimeException` in
30 of 30 runs on each version.
- The `ParquetReadV1Suite` tests with "ids" in their name (7 tests,
including this one) pass locally on the default profile (Spark 4.1) and with
`-Pspark-3.4`. On Spark 3.4 the test now writes two files, of 454 and 476
bytes: the same two-row files as the plain Spark program, and no empty file.
The `run-all-spark-profiles` label is on this PR, so the Comet suites run
against Spark 3.4, 3.5, 4.0 and 4.2 before it lands.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]