laserninja opened a new pull request, #13533: URL: https://github.com/apache/gravitino/pull/13533
### What changes were proposed in this pull request? Add `system_iceberg_rewrite_manifests`, its strategy handler, and the built-in job adapter. Trigger at 500 manifests, or at 100 with average size below 8 MiB, using only a complete measurement for the selected spec. Preserve the collector's resolved spec through evaluation and submission, including after partition evolution. Add boundary, serialization, recommender, and Spark integration tests plus operator/OpenAPI docs. Builds on rewrite-job PR #13329 and the manifest-statistics follow-up (`ea1199562`, imported as `3ac2fecbc`). The policy-only commit is `092f6c119`; prerequisite commits remain separate. ### Why are the changes needed? Connect measured manifest statistics to policy-driven rewrite submission without accidentally evaluating one spec and rewriting another. Related: #11196. Scheduling and cooldowns remain out of scope. ### Does this PR introduce _any_ user-facing change? Adds configurable count/size thresholds, `spec_id`, and `use_caching` to the typed policy. Standalone CLI evaluation requires an explicit spec; collection-cycle callers can retain the collector's resolved default ID. Missing measurements produce no recommendation. ### How was this patch tested? 1,232 tests passed and 6 were skipped across API, common, optimizer API, optimizer, and jobs. Spark 3.5 / Iceberg 1.11.0 integration verifies collection, default-spec evolution, evaluation, adapted submission arguments, and real rewrite execution with unchanged rows and other-spec manifests. Javadocs, `spotlessApply`/checks, `:docs:build`, `rat`, and `git diff --check` passed. Tests use a local Hadoop catalog; remote cluster deployment and an end-to-end REST job submission were not exercised. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
