andygrove opened a new pull request, #6307:
URL: https://github.com/apache/datafusion-comet/pull/6307

   Backport of #6155 and #6111 to `branch-1.1`, as two commits.
   
   Both were cherry-picked in merge order, without conflicts, from 
`14f0f59f74dbcf273ff3cbdb1721a5ee535bbe84` and 
`634e37d08b02da4f9fcbe9dea3606dfd884644c7`. #6111 does not apply on its own, 
because its imports and helper sit on top of #6155's changes to 
`CometIcebergWriteActionSuite`. That is why the two share a pull request.
   
   Every changed line is identical to upstream:
   
   - #6111's diff is byte-identical.
   - Of #6155's 14 files, 11 are identical on `branch-1.1` and on `main` just 
before #6155.
   - The other three differ only in context from later commits that this branch 
also has. In `ci.yml`, #6218's steps move one hunk offset. In `CometConf.scala` 
and `Plugins.scala`, the difference is #6248's deprecation wording.
   
   ## Which issue does this PR close?
   
   Closes #6148 on `branch-1.1`. #6155 already closed it on `main`. #6111 
covers part of #5646.
   
   ## Rationale for this change
   
   Since #6284, a code pull request against `branch-1.1` runs the Iceberg Spark 
test jobs with the native writer enabled, and so will the dispatch before each 
release candidate. As #6148 describes, a passing job can't tell a native write 
from a silent fallback. With #6155, each job's summary page counts the writes 
that ran natively, through the JVM writer, or through Spark's own V2 write, and 
lists why Comet declined. So the release branch's Iceberg runs will show how 
much of 1.1.0's native writer they exercised.
   
   #6111 adds a test for a failed native write attempt. It checks that the task 
retries and that the failed attempt leaves no committed or orphaned data files.
   
   Neither changes what a user's job does. The listener is registered only when 
the internal `spark.comet.testing.icebergWriteReport.dir` is set.
   
   ## What changes are included in this PR?
   
   The original changes, so see #6155 and #6111 for the details. No adaptations 
were needed.
   
   #6155 adds:
   
   - The internal test config `spark.comet.testing.icebergWriteReport.dir`, 
which defaults to `COMET_ICEBERG_WRITE_REPORT_DIR`.
   - `IcebergWriteReportListener`, and the driver plugin hook that registers it.
   - `IcebergTableAsSelectShim`, which recognizes Spark 3.4's CTAS and RTAS 
execs.
   - `dev/ci/summarize-iceberg-writes.py`, and its test, which runs in 
Preflight.
   - Steps in the Iceberg workflow and `dev/local-ci.sh`.
   - A section in the contributor guide's Iceberg Spark tests page.
   
   #6111 adds the test `native acceleration: a mid-write failure retries 
without orphan files` to `CometIcebergWriteActionSuite`, plus a parquet-file 
walker shared by the suite and the test. The suite now runs on `local[5,2]` so 
that a task can retry.
   
   The new config is `.internal()`, so `configs.md` needs no row.
   
   ## How are these changes tested?
   
   Run locally on `branch-1.1` with the default profile (Spark 4.1, Scala 2.13, 
JDK 17):
   
   - `CometIcebergWriteActionSuite` passes, 78 tests, including #6155's three 
report tests and #6111's retry test. `CometPluginsSuite` passes, 6 tests.
   - `dev/ci/test-summarize-iceberg-writes.py` (6 tests), 
`dev/ci/check-ci-config.py` and `dev/ci/check-suites.py` pass.
   - `actionlint` reports the same 26 findings with and without this PR, and 
none of them is in the two changed workflows.
   - `prettier --check` passes on the changed guide.
   
   Upstream, #6155's report tests also passed on Spark 3.4 and 3.5. Against 
`branch-1.1`, the changed paths route this pull request to every suite except 
Spark 3.4's SQL job and the benchmark check. That includes every Spark profile, 
macOS, PyArrow and all four Iceberg versions. The Iceberg jobs should also 
publish this branch's first write report.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to