Hi Shahar and everyone
Jarek, Ash and I were having a small discussion over the community Slack on how 
we can run more CI checks faster and potentially use the merge queue for 
stronger validation before changes reach main.
The main concern is that running the full suite for every queued PR would 
significantly reduce merge throughput with the current runner capacity.
We felt AIP-118 could help make this more practical by moving CI back to larger 
and faster machines.
Apart from RunsOn, two more options for hosted runners came up during our 
discussion:

  1.
BlackSmith <https://www.blacksmith.sh/>  (I know of a couple of popular Open 
Source projects like Chroma, Celery, Turso and Daytona that have adapted this)
  2.
Incredibuild <https://www.incredibuild.com/> (which Ash mentioned was also a 
sponsor at the Airflow summit)

Would love to know what others think, and potentially expedite the proposal.

On 2026/09/18 08:51:34 Shahar Epstein wrote:
> Hello everyone,
>
> Activity in Apache Airflow has skyrocketed over the past year (thanks to
> both humans and their carbon-wasting assistants) and the load on our
> GitHub-hosted CI runners has increased accordingly.
> More often than not, we experience slowness and hiccups with the
> GitHub-hosted runners, primarily because of the number of concurrent jobs.
> This is especially frustrating for maintainers when scheduled canary runs
> need to be restarted or when a PR addressing a high-priority issue, such as
> a security issue, needs to be validated quickly.
> Furthermore, because we share the GitHub-hosted runner capacity with the
> entire ASF organization, Airflow’s usage may also affect other Apache
> projects.
>
> While we continue exploring ways to reduce our CI footprint through code
> and workflow optimizations, we have also investigated offloading some jobs
> to self-hosted runners in cloud environments. This idea has been discussed
> on the dev list several times over the years ([1]
> <https://lists.apache.org/thread/8htrdgf2h8qz1hv7mbb96v8l8x8d1dyl>, [2]
> <https://lists.apache.org/thread/lsnpdovfpnj81pwdlhk768bv8off4nd1>, [3]
> <https://lists.apache.org/thread/55z686pt7377wt21yqjqj26y9s81zcjo>, [4]
> <https://lists.apache.org/thread/ogwjy38hyly9tksfzl294bbp434jokf3>, [5]
> <https://lists.apache.org/thread/4okht98xl127mc9nynpyzxrjdrvj8h0c>), and
> several proofs of concept have been developed. However, it has not yet
> materialized into a permanent solution, partly because of competing
> priorities and partly because optimizations made at the time were
> sufficient.
>
> The problem has continued to grow, so I have decided to tackle it again.
>
> For those who are not aware, Apache Airflow participates in the AWS Open
> Source Credits Program, through which we receive credits once in two years.
> These credits are currently used primarily to host our documentation on S3.
> We recently asked AWS for additional credits to support the self-hosted
> runner effort, and they generously agreed to contribute them.
>
> AIP-118 <https://cwiki.apache.org/confluence/x/-JXwGg> describes the
> motivation for using self-hosted CI runners, the proposed policy governing
> their use, and the available implementation alternatives. I have also
> included an appendix
> <https://cwiki.apache.org/confluence/spaces/AIRFLOW/pages/451974672/AIP-118+Appendix+%E2%80%94+evidence+and+implementation+details?src=contextnavpagetreemode>
> containing supporting data and implementation details.
>
> In general, self-hosted runners would be limited to trusted runs: scheduled
> workflows and runs explicitly approved by committers.
>
> Because of the current budget constraints, we would initially offload
> scheduled runs and selected runs manually labeled by committers. If this
> proves successful, we may ask AWS for additional credits in the future,
> allowing all committer-triggered runs to use self-hosted runners by default.
>
> So far, I have tested two implementation approaches:
>
> *- EKS:* A Kubernetes cluster using Spot Instances. This is cost-effective
> but requires us to operate and maintain the cluster.
>
> *- CodeBuild-managed GitHub Actions runners:* This requires significantly
> less maintenance, but it is x7 times as expensive as the EKS option. I have
> therefore ruled it out for now.
> I am also evaluating a 3rd option: RunsOn <https://runs-on.com/>. It
> appears to provide many of the advantages of the EKS approach with a much
> simpler infrastructure setup and may cost less. However, it would introduce
> reliance on a third party and requires public acknowledgement of its use.
>
> AIP-118 does not propose any changes affecting Apache Airflow users.
> Nevertheless, because it would significantly affect the contribution
> workflows, I believe it is appropriate to put it to a formal vote. I would
> be happy to hear any concerns, suggestions, or alternative ideas in this
> thread before going into a vote.
>
> I'd like to thank:
> - Hussein, Jarek, and Ash for their previous experience with self-hosted
> runners, which was invaluable in formulating the AIP and developing the EKS
> design.
> - Niko for his tremendous help in securing the additional AWS credits.
>
>
> Shahar
>

Reply via email to