Thanks for the feedback!

I agree that using a merge queue could be very beneficial, however - I'd
leave it out of the scope of the current discussion, it can be discussed
and implemented independently.

Regarding the other alternatives, thanks for pointing out - I'll try to
contact their respresentatives to see if and how we could utilize their
services. As commented on the AIP, having redundancy in this case is very
helpful, as being dependent on single 3rd parties for the CI is quite risky
in the long term (this is one of the most critical things in the open
source).

---

At this point of time I plan to proceed with the offerring by runs-on.com,
for the following reasonings:
1. The option for Open Source licensing is publically available on their
docs.
2. I've tested and managed to deploy their stack succesfully and smoothly.
3. Other Apache projects already utilize their services (including Arrow,
Datafusion, and ASF), which means a way easier integration.
4. I met up with runs-on's CEO on a video call, Mr. Cyril Rohr, and he
expressed his willigness to support the onboarding of Apache Airflow as
needed.

I updated the AIP + appendix accordingly. Please note that:
1. The AIP is more focused on the "why we need it" and "what alternatives
exist".
2. If at later point we need to alter the implementation for any reason
(e.g., not enough credits on AWS / runs-on.com stops their free license
support / CI gets even more complex) - we always have the option to use the
alternatives (e.g., Azure / custom EKS). As I previously mentioned, I
support in exploring other alternatives - but it doesn't have to block this
AIP or its implementation.

I will wait for the costs report arrives tomorrow, and assuming that there
won't be any surprises or objections -
I'll set up a vote during the next week.


Shahar

On Wed, Sep 23, 2026 at 7:18 PM Subhramit B.B. <[email protected]> wrote:

> Hi Shahar and everyone
> Jarek, Ash and I were having a small discussion over the community Slack
> on how we can run more CI checks faster and potentially use the merge queue
> for stronger validation before changes reach main.
> The main concern is that running the full suite for every queued PR would
> significantly reduce merge throughput with the current runner capacity.
> We felt AIP-118 could help make this more practical by moving CI back to
> larger and faster machines.
> Apart from RunsOn, two more options for hosted runners came up during our
> discussion:
>
>   1.
> BlackSmith <https://www.blacksmith.sh/>  (I know of a couple of popular
> Open Source projects like Chroma, Celery, Turso and Daytona that have
> adapted this)
>   2.
> Incredibuild <https://www.incredibuild.com/> (which Ash mentioned was
> also a sponsor at the Airflow summit)
>
> Would love to know what others think, and potentially expedite the
> proposal.
>
> On 2026/09/18 08:51:34 Shahar Epstein wrote:
> > Hello everyone,
> >
> > Activity in Apache Airflow has skyrocketed over the past year (thanks to
> > both humans and their carbon-wasting assistants) and the load on our
> > GitHub-hosted CI runners has increased accordingly.
> > More often than not, we experience slowness and hiccups with the
> > GitHub-hosted runners, primarily because of the number of concurrent
> jobs.
> > This is especially frustrating for maintainers when scheduled canary runs
> > need to be restarted or when a PR addressing a high-priority issue, such
> as
> > a security issue, needs to be validated quickly.
> > Furthermore, because we share the GitHub-hosted runner capacity with the
> > entire ASF organization, Airflow’s usage may also affect other Apache
> > projects.
> >
> > While we continue exploring ways to reduce our CI footprint through code
> > and workflow optimizations, we have also investigated offloading some
> jobs
> > to self-hosted runners in cloud environments. This idea has been
> discussed
> > on the dev list several times over the years ([1]
> > <https://lists.apache.org/thread/8htrdgf2h8qz1hv7mbb96v8l8x8d1dyl>, [2]
> > <https://lists.apache.org/thread/lsnpdovfpnj81pwdlhk768bv8off4nd1>, [3]
> > <https://lists.apache.org/thread/55z686pt7377wt21yqjqj26y9s81zcjo>, [4]
> > <https://lists.apache.org/thread/ogwjy38hyly9tksfzl294bbp434jokf3>, [5]
> > <https://lists.apache.org/thread/4okht98xl127mc9nynpyzxrjdrvj8h0c>), and
> > several proofs of concept have been developed. However, it has not yet
> > materialized into a permanent solution, partly because of competing
> > priorities and partly because optimizations made at the time were
> > sufficient.
> >
> > The problem has continued to grow, so I have decided to tackle it again.
> >
> > For those who are not aware, Apache Airflow participates in the AWS Open
> > Source Credits Program, through which we receive credits once in two
> years.
> > These credits are currently used primarily to host our documentation on
> S3.
> > We recently asked AWS for additional credits to support the self-hosted
> > runner effort, and they generously agreed to contribute them.
> >
> > AIP-118 <https://cwiki.apache.org/confluence/x/-JXwGg> describes the
> > motivation for using self-hosted CI runners, the proposed policy
> governing
> > their use, and the available implementation alternatives. I have also
> > included an appendix
> > <
> https://cwiki.apache.org/confluence/spaces/AIRFLOW/pages/451974672/AIP-118+Appendix+%E2%80%94+evidence+and+implementation+details?src=contextnavpagetreemode
> >
> > containing supporting data and implementation details.
> >
> > In general, self-hosted runners would be limited to trusted runs:
> scheduled
> > workflows and runs explicitly approved by committers.
> >
> > Because of the current budget constraints, we would initially offload
> > scheduled runs and selected runs manually labeled by committers. If this
> > proves successful, we may ask AWS for additional credits in the future,
> > allowing all committer-triggered runs to use self-hosted runners by
> default.
> >
> > So far, I have tested two implementation approaches:
> >
> > *- EKS:* A Kubernetes cluster using Spot Instances. This is
> cost-effective
> > but requires us to operate and maintain the cluster.
> >
> > *- CodeBuild-managed GitHub Actions runners:* This requires significantly
> > less maintenance, but it is x7 times as expensive as the EKS option. I
> have
> > therefore ruled it out for now.
> > I am also evaluating a 3rd option: RunsOn <https://runs-on.com/>. It
> > appears to provide many of the advantages of the EKS approach with a much
> > simpler infrastructure setup and may cost less. However, it would
> introduce
> > reliance on a third party and requires public acknowledgement of its use.
> >
> > AIP-118 does not propose any changes affecting Apache Airflow users.
> > Nevertheless, because it would significantly affect the contribution
> > workflows, I believe it is appropriate to put it to a formal vote. I
> would
> > be happy to hear any concerns, suggestions, or alternative ideas in this
> > thread before going into a vote.
> >
> > I'd like to thank:
> > - Hussein, Jarek, and Ash for their previous experience with self-hosted
> > runners, which was invaluable in formulating the AIP and developing the
> EKS
> > design.
> > - Niko for his tremendous help in securing the additional AWS credits.
> >
> >
> > Shahar
> >
>

Reply via email to