andygrove commented on issue #6301:
URL:
https://github.com/apache/datafusion-comet/issues/6301#issuecomment-5872795614
This is a flaky test rather than a regression. The only failure is
`ParquetReadV1Suite` "a file without ids next to a file with ids is checked on
its own", which #6116 added. The assertion that failed was on the Spark
reference error, not Comet's. The same test passed in the Spark 3.4 `[scans]`
job of the 2026-09-27 nightly, which already included #6116.
**Cause.** The test writes each side with `sparkContext.parallelize` on the
`local[5]` test session. Spark always writes a file for partition 0, so each
write leaves an empty `part-00000` next to two one-row files. The read packs
the six files into three tasks, and on Spark 3.4 and 3.5 two of them fail with
different exceptions:
| Task | Files | Exception |
| --- | --- | --- |
| 0 | two files with ids | none |
| 1 | two files without ids | `RuntimeException` from `ParquetReadSupport` |
| 2 | empty file with ids, then empty file without ids |
`SparkException("Encountered error while reading file")` wrapping the
`RuntimeException` |
`FileScanRDD.nextIterator` adds that wrapper only when the error is raised
inside its `try { hasNext }`. A file without ids only fails there when it
follows an empty file in the same task. The 3.x `DAGScheduler` wraps whichever
task fails first in "Job aborted", and `isMissingFieldIdsError` looks only at
`getCause`. The test passes when task 1 fails first and fails when task 2 does.
This run reported `Task 2 in stage 2468.0 failed`.
- Spark 3.5 has the same exposure; it just has not lost the race in CI yet.
- Spark 4.x wraps every read error once in `FAILED_READ_FILE` and passes it
through without "Job aborted", so the test is deterministic there.
- Comet raises the same exception from every task on every version.
I checked this with a plain Spark program (no Comet) that repeats the test's
writes and drains each read partition on its own. It shows the layout and both
exception shapes above on Spark 3.4.3 and 3.5.9.
**Fix.** #6312 writes each side as one file. With no empty file, only one
task fails, and Spark raises the bare `RuntimeException` every time.
`branch-1.1` has the same test through #6266 and will need the same change.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]