Hi all, A few of us have been working on this proposal and would like community feedback on:
AIP-97 Task Failure Classification https://cwiki.apache.org/confluence/x/0ITMFw Airflow’s executor runs separately from the task worker, so it can still read backend failure details after the worker stops. AIP-97 passes confirmed causes, such as Kubernetes preemption, into failure handling for listeners, logs and metrics. When the cause is unclear, it stays unclassified. This helps integrations distinguish infrastructure interruptions from application errors and route alerts appropriately. Ordinary retries, clears and DAG callback interfaces stay unchanged. We’ve separated infrastructure retry decisions into: AIP-122 Infrastructure Retries https://cwiki.apache.org/confluence/x/-pXwGg When enabled, AIP-122 uses confirmed infrastructure failures to preserve ordinary task retries, within an administrator-configured limit. For example, consider a Spark task with retries=1. Its first attempt is interrupted by a confirmed Kubernetes preemption. The retry starts, but the Spark job fails during that second attempt. There is now no retry left: the user-defined retry was consumed by an infrastructure interruption rather than an application failure. With AIP-122 enabled, the first infrastructure failure can receive an infrastructure retry, leaving the ordinary retry available for the later Spark failure. AIP-97 can deliver value independently while we work with the community on AIP-122. Both AIPs include implementation POCs, runnable examples and end-to-end evidence. Thank you, Stefan, on behalf of the AIP-97 and AIP-122 co-authors Reference: Previous dev-list discussion <https://lists.apache.org/thread/g4jhd0vz71x6jm4z8p6hjv9z7rmtb3jk>, bcc’d folks from there FYI.
