Sounds cool.

On Fri, Sep 25, 2026 at 3:13 PM Shahar Epstein <[email protected]> wrote:

> Thanks for the feedback!
>
> I agree that using a merge queue could be very beneficial, however - I'd
> leave it out of the scope of the current discussion, it can be discussed
> and implemented independently.
>
> Regarding the other alternatives, thanks for pointing out - I'll try to
> contact their respresentatives to see if and how we could utilize their
> services. As commented on the AIP, having redundancy in this case is very
> helpful, as being dependent on single 3rd parties for the CI is quite risky
> in the long term (this is one of the most critical things in the open
> source).
>
> ---
>
> At this point of time I plan to proceed with the offerring by runs-on.com,
> for the following reasonings:
> 1. The option for Open Source licensing is publically available on their
> docs.
> 2. I've tested and managed to deploy their stack succesfully and smoothly.
> 3. Other Apache projects already utilize their services (including Arrow,
> Datafusion, and ASF), which means a way easier integration.
> 4. I met up with runs-on's CEO on a video call, Mr. Cyril Rohr, and he
> expressed his willigness to support the onboarding of Apache Airflow as
> needed.
>
> I updated the AIP + appendix accordingly. Please note that:
> 1. The AIP is more focused on the "why we need it" and "what alternatives
> exist".
> 2. If at later point we need to alter the implementation for any reason
> (e.g., not enough credits on AWS / runs-on.com stops their free license
> support / CI gets even more complex) - we always have the option to use the
> alternatives (e.g., Azure / custom EKS). As I previously mentioned, I
> support in exploring other alternatives - but it doesn't have to block this
> AIP or its implementation.
>
> I will wait for the costs report arrives tomorrow, and assuming that there
> won't be any surprises or objections -
> I'll set up a vote during the next week.
>
>
> Shahar
>
> On Wed, Sep 23, 2026 at 7:18 PM Subhramit B.B. <[email protected]>
> wrote:
>
> > Hi Shahar and everyone
> > Jarek, Ash and I were having a small discussion over the community Slack
> > on how we can run more CI checks faster and potentially use the merge
> queue
> > for stronger validation before changes reach main.
> > The main concern is that running the full suite for every queued PR would
> > significantly reduce merge throughput with the current runner capacity.
> > We felt AIP-118 could help make this more practical by moving CI back to
> > larger and faster machines.
> > Apart from RunsOn, two more options for hosted runners came up during our
> > discussion:
> >
> >   1.
> > BlackSmith <https://www.blacksmith.sh/>  (I know of a couple of popular
> > Open Source projects like Chroma, Celery, Turso and Daytona that have
> > adapted this)
> >   2.
> > Incredibuild <https://www.incredibuild.com/> (which Ash mentioned was
> > also a sponsor at the Airflow summit)
> >
> > Would love to know what others think, and potentially expedite the
> > proposal.
> >
> > On 2026/09/18 08:51:34 Shahar Epstein wrote:
> > > Hello everyone,
> > >
> > > Activity in Apache Airflow has skyrocketed over the past year (thanks
> to
> > > both humans and their carbon-wasting assistants) and the load on our
> > > GitHub-hosted CI runners has increased accordingly.
> > > More often than not, we experience slowness and hiccups with the
> > > GitHub-hosted runners, primarily because of the number of concurrent
> > jobs.
> > > This is especially frustrating for maintainers when scheduled canary
> runs
> > > need to be restarted or when a PR addressing a high-priority issue,
> such
> > as
> > > a security issue, needs to be validated quickly.
> > > Furthermore, because we share the GitHub-hosted runner capacity with
> the
> > > entire ASF organization, Airflow’s usage may also affect other Apache
> > > projects.
> > >
> > > While we continue exploring ways to reduce our CI footprint through
> code
> > > and workflow optimizations, we have also investigated offloading some
> > jobs
> > > to self-hosted runners in cloud environments. This idea has been
> > discussed
> > > on the dev list several times over the years ([1]
> > > <https://lists.apache.org/thread/8htrdgf2h8qz1hv7mbb96v8l8x8d1dyl>,
> [2]
> > > <https://lists.apache.org/thread/lsnpdovfpnj81pwdlhk768bv8off4nd1>,
> [3]
> > > <https://lists.apache.org/thread/55z686pt7377wt21yqjqj26y9s81zcjo>,
> [4]
> > > <https://lists.apache.org/thread/ogwjy38hyly9tksfzl294bbp434jokf3>,
> [5]
> > > <https://lists.apache.org/thread/4okht98xl127mc9nynpyzxrjdrvj8h0c>),
> and
> > > several proofs of concept have been developed. However, it has not yet
> > > materialized into a permanent solution, partly because of competing
> > > priorities and partly because optimizations made at the time were
> > > sufficient.
> > >
> > > The problem has continued to grow, so I have decided to tackle it
> again.
> > >
> > > For those who are not aware, Apache Airflow participates in the AWS
> Open
> > > Source Credits Program, through which we receive credits once in two
> > years.
> > > These credits are currently used primarily to host our documentation on
> > S3.
> > > We recently asked AWS for additional credits to support the self-hosted
> > > runner effort, and they generously agreed to contribute them.
> > >
> > > AIP-118 <https://cwiki.apache.org/confluence/x/-JXwGg> describes the
> > > motivation for using self-hosted CI runners, the proposed policy
> > governing
> > > their use, and the available implementation alternatives. I have also
> > > included an appendix
> > > <
> >
> https://cwiki.apache.org/confluence/spaces/AIRFLOW/pages/451974672/AIP-118+Appendix+%E2%80%94+evidence+and+implementation+details?src=contextnavpagetreemode
> > >
> > > containing supporting data and implementation details.
> > >
> > > In general, self-hosted runners would be limited to trusted runs:
> > scheduled
> > > workflows and runs explicitly approved by committers.
> > >
> > > Because of the current budget constraints, we would initially offload
> > > scheduled runs and selected runs manually labeled by committers. If
> this
> > > proves successful, we may ask AWS for additional credits in the future,
> > > allowing all committer-triggered runs to use self-hosted runners by
> > default.
> > >
> > > So far, I have tested two implementation approaches:
> > >
> > > *- EKS:* A Kubernetes cluster using Spot Instances. This is
> > cost-effective
> > > but requires us to operate and maintain the cluster.
> > >
> > > *- CodeBuild-managed GitHub Actions runners:* This requires
> significantly
> > > less maintenance, but it is x7 times as expensive as the EKS option. I
> > have
> > > therefore ruled it out for now.
> > > I am also evaluating a 3rd option: RunsOn <https://runs-on.com/>. It
> > > appears to provide many of the advantages of the EKS approach with a
> much
> > > simpler infrastructure setup and may cost less. However, it would
> > introduce
> > > reliance on a third party and requires public acknowledgement of its
> use.
> > >
> > > AIP-118 does not propose any changes affecting Apache Airflow users.
> > > Nevertheless, because it would significantly affect the contribution
> > > workflows, I believe it is appropriate to put it to a formal vote. I
> > would
> > > be happy to hear any concerns, suggestions, or alternative ideas in
> this
> > > thread before going into a vote.
> > >
> > > I'd like to thank:
> > > - Hussein, Jarek, and Ash for their previous experience with
> self-hosted
> > > runners, which was invaluable in formulating the AIP and developing the
> > EKS
> > > design.
> > > - Niko for his tremendous help in securing the additional AWS credits.
> > >
> > >
> > > Shahar
> > >
> >
>

Reply via email to