Hi Jung-Hyun Andrew Kim and Teghveer Singh Ateliey, Thanks for the doc. I went through it and it looks good to me. If there is no other feedbacks/concerns, I think it is safe to move to implementation.
Vincent On 2026/07/13 03:40:53 Jung-Hyun Andrew Kim wrote: > Hi Everyone, > Following up on our previous discussion regarding the /api/v2/monitor/health > endpoint, Teghveer Singh Ateliey and I went ahead and worked through all your > feedback together and made a full proposal doc as well as an example > implementation for your review! > The Proposal > Teghveer and I propose that we update the endpoint to return both a detailed > status field and instances field so as to report an aggregated status of the > instance's health, and of the individual instances. We would retain the old > fields to remain backwards compatible. > Stretch Goal: Extend this coverage to include Executor health reporting, > catching edge-case failure modes within individual schedulers that currently > fail silently. > Resources & Next Steps > We have put together a detailed design doc and a draft implementation to show > how this would look in practice: > Design Document: Google Doc Link > <https://docs.google.com/document/d/1yWO-0YL6TNpABLwXoKy3YGVFZGnr4CE0Q1iZ_rpNmBo/edit?usp=sharing> > Example Implementation: GitHub Pull Request > <https://github.com/JH-A-Kim/airflow-fork/pull/4> > Please take a look and share your thoughts on the approach, the proposed JSON > structure changes, or any edge cases we should consider. > Looking forward to your feedback! > Best, > Jung-Hyun Andrew Kim and Teghveer Singh Ateliey > On 2026/06/25 00:01:00 Jung-Hyun Kim wrote: > > The Problem > > In distributed Airflow environments running multiple schedulers, the > > current health endpoint contains a significant monitoring blind spot. > > Currently, the health check determines the status of the scheduler by > > querying the metadata database using the most_recent_job method found in > > job.py: > > > > @provide_session > > def most_recent_job(job_type: str, *, session: Session = NEW_SESSION) -> > > Job | None: > > """ > > Return the most recent job of this type, if any, based on last > > heartbeat received. > > > > Jobs in "running" state take precedence over others to make sure alive > > job is returned if it is available. > > > > :param job_type: job type to query for to get the most recent job for > > :param session: Database session > > :end_date: None > > """ > > return session.scalar( > > select(Job) > > .where(Job.job_type == job_type) > > .order_by( > > # Put "running" jobs at the front. > > case({JobState.RUNNING: 0}, value=Job.state, else_=1), > > Job.latest_heartbeat.desc(), > > ) > > .limit(1) > > ) > > > > > > This database query explicitly sorts records by the RUNNING state and > > applies .limit(1), returning only a single, absolute newest job record. > > This result is then processed in airflow_health.py via the > > get_airflow_health endpoint method: > > > > def get_airflow_health() -> dict[str, Any]: > > """Get the health for Airflow metadatabase, scheduler and triggerer.""" > > metadatabase_status = HEALTHY > > latest_scheduler_heartbeat = None > > latest_triggerer_heartbeat = None > > latest_dag_processor_heartbeat = None > > > > scheduler_status = UNHEALTHY > > triggerer_status: str | None = UNHEALTHY > > dag_processor_status: str | None = UNHEALTHY > > > > try: > > latest_scheduler_job = SchedulerJobRunner.most_recent_job() > > > > if latest_scheduler_job: > > if latest_scheduler_job.latest_heartbeat: > > latest_scheduler_heartbeat = > > latest_scheduler_job.latest_heartbeat.isoformat() > > if latest_scheduler_job.is_alive(): > > scheduler_status = HEALTHY > > except Exception: > > metadatabase_status = UNHEALTHY > > > > > > Because the health endpoint evaluates only the single job returned by > > most_recent_job(), the check can only ever validate the health of one > > scheduler at a time. > > In a distributed deployment with multiple active schedulers, if even one > > instance is running cleanly, the endpoint will flag as healthy even if all > > other parallel scheduler instances have gone down. > > To get meaningful information regarding the scheduler status from the > > health endpoint it is worth it to monitor every scheduler in the > > distributed environment instead of just a single scheduler. > > The Proposed Solution > > To deal with this problem we can add a new field called schedulers (plural > > for multiple schedulers) in the health endpoint that returns a 3-tier > > aggregated status that covers the following: > > > > * > > HEALTHY: All registered scheduler instances are fully operational and > > actively heartbeating. > > * > > DEGRADED: At least one scheduler instance is down or failing, but at least > > one remaining instance is still working. > > * > > DOWN: All scheduler instances have failed or stopped working. > > > > Per-Instance Diagnostic Breakdown > > We should also add a per instance breakdown as a nested list that will show > > the following: > > > > 1. > > hostname > > 2. > > status: Individual status > > 3. > > latest_heartbeat > > > > Example > > > > { > > "metadatabase": { > > "status": "healthy" > > }, > > "scheduler": { > > "scheduler_status": "healthy", > > "latest_scheduler_heartbeat": "2026-06-24T23:15:02+00:00" > > }, > > "schedulers": { > > "status": "DEGRADED", > > "instances": [ > > { > > "hostname": "scheduler-ha-instance-1", > > "status": "HEALTHY", > > "latest_heartbeat": "2026-06-24T23:15:02+00:00" > > }, > > { > > "hostname": "scheduler-ha-instance-2", > > "status": "DOWN", > > "latest_heartbeat": "2026-06-24T23:10:14+00:00" > > }, > > { > > "hostname": "scheduler-ha-instance-3", > > "status": "HEALTHY", > > "latest_heartbeat": "2026-06-24T23:14:59+00:00" > > } > > ] > > } > > } > > > > Could end up looking roughly like this, resulting in a more meaningful > > health endpoint that will make it easier to diagnose issues with the > > scheduler. This is a LAZY CONSENSUS proposal. > > --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
