danny0405 opened a new issue, #20017: URL: https://github.com/apache/hudi/issues/20017
### Bug Description **What happened:** A binary copy can fail after some input row groups have been copied. Failure cleanup calls the copier's existing `close()` method, which finalizes metadata and writes the Parquet footer. When that finalization succeeds, an incomplete copy remains as a readable Parquet file. `HoodieBinaryCopyHandle` does not create a write marker, and `SparkBinaryCopyClusteringExecutionStrategy` chooses a new file ID for each invocation. A retried task can therefore leave its previous attempt's output behind. This is tracked separately from resource release in [PR #18776](https://github.com/apache/hudi/pull/18776#discussion_r4060375253). **What you expected:** An unsuccessful binary-copy attempt should have a defined cleanup path for its output. Investigate abort/delete semantics or integration with marker/rollback cleanup without deleting a successful retry's output. **Steps to reproduce:** 1. Binary-copy multiple input Parquet files into one new output file. 2. Let the first input complete, then inject an IOException while fetching the next input. 3. Close the copier to release resources and inspect the partial output. 4. Retry the clustering task with a fresh output file ID and inspect what happens to the first attempt's file during commit, rollback, and later cleaning. The PR's `testCloseAfterPrefetchFailureReleasesResources` exercises the first-input-success/second-input-failure path. A dedicated regression should additionally verify orphan removal across retry and cleanup. **Why regular cleaner is not a guarantee:** - [CleanPlanner](https://github.com/apache/hudi/blob/f8affff7793984a7a323988ac54598155842364f/hudi-client/hudi-client-common/src/main/java/org/apache/hudi/table/action/clean/CleanPlanner.java) operates on file-group slices and retains the latest slice under the usual retention policies. An orphan with its own file ID may be that group's only slice. - [HoodieFileGroup](https://github.com/apache/hudi/blob/f8affff7793984a7a323988ac54598155842364f/hudi-common/src/main/java/org/apache/hudi/common/model/HoodieFileGroup.java) filters `getAllFileSlices()` through the commit timeline, so inflight/uncommitted output is not simply treated as old committed data to clean. - Listing-based rollback can remove files associated with a rolled-back instant, but that is different from the cleaner, and a successful task retry does not necessarily roll back the whole clustering instant. These are code-traced reasons not to rely on regular cleaner as the orphan-recovery mechanism; cleanup behavior across metadata-table and rollback modes needs dedicated validation. ### Environment **Hudi version:** 1.3.0-SNAPSHOT, PR #18776 at `f8affff77939`. **Query engine:** Spark clustering using `SparkBinaryCopyClusteringExecutionStrategy`. **Relevant configs:** Binary-copy clustering enabled and an input read/prefetch failure after partial progress. ### Logs and Stack Trace The injected failure is `IOException: Failed to retrieve prefetched content ...`. The current PR regression verifies resource release; it does not claim to verify end-to-end orphan recovery. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
