Sounds cool. On Fri, Sep 25, 2026 at 3:13 PM Shahar Epstein <[email protected]> wrote:
> Thanks for the feedback! > > I agree that using a merge queue could be very beneficial, however - I'd > leave it out of the scope of the current discussion, it can be discussed > and implemented independently. > > Regarding the other alternatives, thanks for pointing out - I'll try to > contact their respresentatives to see if and how we could utilize their > services. As commented on the AIP, having redundancy in this case is very > helpful, as being dependent on single 3rd parties for the CI is quite risky > in the long term (this is one of the most critical things in the open > source). > > --- > > At this point of time I plan to proceed with the offerring by runs-on.com, > for the following reasonings: > 1. The option for Open Source licensing is publically available on their > docs. > 2. I've tested and managed to deploy their stack succesfully and smoothly. > 3. Other Apache projects already utilize their services (including Arrow, > Datafusion, and ASF), which means a way easier integration. > 4. I met up with runs-on's CEO on a video call, Mr. Cyril Rohr, and he > expressed his willigness to support the onboarding of Apache Airflow as > needed. > > I updated the AIP + appendix accordingly. Please note that: > 1. The AIP is more focused on the "why we need it" and "what alternatives > exist". > 2. If at later point we need to alter the implementation for any reason > (e.g., not enough credits on AWS / runs-on.com stops their free license > support / CI gets even more complex) - we always have the option to use the > alternatives (e.g., Azure / custom EKS). As I previously mentioned, I > support in exploring other alternatives - but it doesn't have to block this > AIP or its implementation. > > I will wait for the costs report arrives tomorrow, and assuming that there > won't be any surprises or objections - > I'll set up a vote during the next week. > > > Shahar > > On Wed, Sep 23, 2026 at 7:18 PM Subhramit B.B. <[email protected]> > wrote: > > > Hi Shahar and everyone > > Jarek, Ash and I were having a small discussion over the community Slack > > on how we can run more CI checks faster and potentially use the merge > queue > > for stronger validation before changes reach main. > > The main concern is that running the full suite for every queued PR would > > significantly reduce merge throughput with the current runner capacity. > > We felt AIP-118 could help make this more practical by moving CI back to > > larger and faster machines. > > Apart from RunsOn, two more options for hosted runners came up during our > > discussion: > > > > 1. > > BlackSmith <https://www.blacksmith.sh/> (I know of a couple of popular > > Open Source projects like Chroma, Celery, Turso and Daytona that have > > adapted this) > > 2. > > Incredibuild <https://www.incredibuild.com/> (which Ash mentioned was > > also a sponsor at the Airflow summit) > > > > Would love to know what others think, and potentially expedite the > > proposal. > > > > On 2026/09/18 08:51:34 Shahar Epstein wrote: > > > Hello everyone, > > > > > > Activity in Apache Airflow has skyrocketed over the past year (thanks > to > > > both humans and their carbon-wasting assistants) and the load on our > > > GitHub-hosted CI runners has increased accordingly. > > > More often than not, we experience slowness and hiccups with the > > > GitHub-hosted runners, primarily because of the number of concurrent > > jobs. > > > This is especially frustrating for maintainers when scheduled canary > runs > > > need to be restarted or when a PR addressing a high-priority issue, > such > > as > > > a security issue, needs to be validated quickly. > > > Furthermore, because we share the GitHub-hosted runner capacity with > the > > > entire ASF organization, Airflow’s usage may also affect other Apache > > > projects. > > > > > > While we continue exploring ways to reduce our CI footprint through > code > > > and workflow optimizations, we have also investigated offloading some > > jobs > > > to self-hosted runners in cloud environments. This idea has been > > discussed > > > on the dev list several times over the years ([1] > > > <https://lists.apache.org/thread/8htrdgf2h8qz1hv7mbb96v8l8x8d1dyl>, > [2] > > > <https://lists.apache.org/thread/lsnpdovfpnj81pwdlhk768bv8off4nd1>, > [3] > > > <https://lists.apache.org/thread/55z686pt7377wt21yqjqj26y9s81zcjo>, > [4] > > > <https://lists.apache.org/thread/ogwjy38hyly9tksfzl294bbp434jokf3>, > [5] > > > <https://lists.apache.org/thread/4okht98xl127mc9nynpyzxrjdrvj8h0c>), > and > > > several proofs of concept have been developed. However, it has not yet > > > materialized into a permanent solution, partly because of competing > > > priorities and partly because optimizations made at the time were > > > sufficient. > > > > > > The problem has continued to grow, so I have decided to tackle it > again. > > > > > > For those who are not aware, Apache Airflow participates in the AWS > Open > > > Source Credits Program, through which we receive credits once in two > > years. > > > These credits are currently used primarily to host our documentation on > > S3. > > > We recently asked AWS for additional credits to support the self-hosted > > > runner effort, and they generously agreed to contribute them. > > > > > > AIP-118 <https://cwiki.apache.org/confluence/x/-JXwGg> describes the > > > motivation for using self-hosted CI runners, the proposed policy > > governing > > > their use, and the available implementation alternatives. I have also > > > included an appendix > > > < > > > https://cwiki.apache.org/confluence/spaces/AIRFLOW/pages/451974672/AIP-118+Appendix+%E2%80%94+evidence+and+implementation+details?src=contextnavpagetreemode > > > > > > containing supporting data and implementation details. > > > > > > In general, self-hosted runners would be limited to trusted runs: > > scheduled > > > workflows and runs explicitly approved by committers. > > > > > > Because of the current budget constraints, we would initially offload > > > scheduled runs and selected runs manually labeled by committers. If > this > > > proves successful, we may ask AWS for additional credits in the future, > > > allowing all committer-triggered runs to use self-hosted runners by > > default. > > > > > > So far, I have tested two implementation approaches: > > > > > > *- EKS:* A Kubernetes cluster using Spot Instances. This is > > cost-effective > > > but requires us to operate and maintain the cluster. > > > > > > *- CodeBuild-managed GitHub Actions runners:* This requires > significantly > > > less maintenance, but it is x7 times as expensive as the EKS option. I > > have > > > therefore ruled it out for now. > > > I am also evaluating a 3rd option: RunsOn <https://runs-on.com/>. It > > > appears to provide many of the advantages of the EKS approach with a > much > > > simpler infrastructure setup and may cost less. However, it would > > introduce > > > reliance on a third party and requires public acknowledgement of its > use. > > > > > > AIP-118 does not propose any changes affecting Apache Airflow users. > > > Nevertheless, because it would significantly affect the contribution > > > workflows, I believe it is appropriate to put it to a formal vote. I > > would > > > be happy to hear any concerns, suggestions, or alternative ideas in > this > > > thread before going into a vote. > > > > > > I'd like to thank: > > > - Hussein, Jarek, and Ash for their previous experience with > self-hosted > > > runners, which was invaluable in formulating the AIP and developing the > > EKS > > > design. > > > - Niko for his tremendous help in securing the additional AWS credits. > > > > > > > > > Shahar > > > > > >
