Yeah. Subscription cannot be used as an API replacement "really" (at least long term and this is how subscriptions are being "targeted"). You really need to use them via the "Claude" or "Codex" CLI. While there are some workarounds (like `claude -p`), this method might stop working (it was scheduled for disabling - but Anthropic backed out temporarily).
But yeah absolutely - we can use the Azure Credits, for that as well. Ideally this will be **very little** credit use. We could even use Edge executors for that (which is even more Dogfooding) - even if Airflow will be running in Amazon - the runners could be set-up in Azure for example. Or we could connect to the models Azure exposes. >From my Magpie experiments about 80%-90% of such checks can be purely deterministic, and the vast majority of the AI usage falls into two cases: * generating responses (most of them can be templated - only small amount of tokens is spent to make messages "tailored" to the PR/ Security issue) * decision making on how to classify things (which can be way cheaper now with Jev - and similar solutions - Open AI just 3 days ago released "Decisions API" - and pretty much everyone will have similar services soon [1] And i think we should treat it as a starting point - there are and will be many sources of tokens we can use: * We have stakeholders - Google and Amazon—who can likely donate some of their credits * ASF is finalizing llm service [2] for projects to use (open models on rented GPUs, plus donated credits to ASF by big models). This is specifically foreseen as a shared resource that ASF projects can use, and using it for "around CI" work is one of the stronger use-cases. I think it's important that we monitor and use the capability of our subscriptions to iterate on ideas with AI - and turn them in **mostly** deterministic workflows with as-small-as-possible usage of tokens and monitor it. This is one of the things we have almost done in Magpie -> - we have a SKILL that optimizes other skills by simplifying language, refactoring and extracting pieces of it - and turning everything that can be turned into deterministict Python code. And I am plannning to use it on Airflow AGENTS.md and SKILLs (see the other discussion). [1] https://community.openai.com/t/decisions-api-is-now-available-in-public-beta/1403877 [2] https://llm.apache.org - LLM service from ASF that is about to be announced [3] Skill optimization effort in Magpie https://github.com/apache/magpie/issues/1342 J. On Fri, Oct 9, 2026 at 12:07 PM Shubham Raj <[email protected]> wrote: > Thanks Ash for the quick response and feedback, > > Initially, yes (while in shadow mode), I plan to use copy/paste or GH APIs. > But the final goal would be like using a service account creds, which will > automatically post the reviews on the PR when approved, Today when we use > the harness to post the comments, it ends with "Drafted-by: <harness> > (<model>); reviewed by <username> before posting", we will follow the same > thing, it's just the username would be the maintainer's name who approved > the HITL. > > Initially, I'm thinking of implementing this for committers+, and we can > decide on scaling later based on guardrails or rules implementation. > > Yes, those are good programs, but they are specific to users. I think that > would be really needed for contributors regarding their own PR creation and > related work. I don't think the program provides API keys for developers, > which we need for the setup. Good to know we have Azure credits; this adds > to the viable options. > > Thanks, > Shubham > > On Fri, Oct 9, 2026 at 4:18 PM Ash Berlin-Taylor <[email protected]> wrote: > > > Love the idea of using Airflow to manage the Airflow project. > > > > Some clarifying questions: > > > > > human-approved comments, labels, draft conversion, closure where > > permitted, and review dispositions; > > > > For the initial trail, and in the long term, who or what is making > posting > > comments? Initial version is “the human has to copy them etc” (and we > could > > make that easier by providing `gh` commands etc) or an “AirflowBot” GH > > user/app etc? (I think I’m fine either way, I’m just unsure) > > > > Who would have access to see the results of these runs? Who would be able > > to be the Human in the HITL to approve things? Were you thinking > everyone > > in Triage, Committers and PMC groups? I think from reading the AIP you > have > > it as Committers+? > > > > Another possible triage rule: history of opening PRs either too quickly, > > or to many and then never responding to comments? Might not be a real > case > > so might not be worth it. > > > > > > On question 4: Both Anthropic and OpenAI have “OSS” programs where they > > give (rolling, after re-applying) 6 month free things to contributors of > > projects like Airflow. Downside to that is it would be tied to an > > individual user I suspect. We also have a large number of Azure credits > > that have a (now 1year left) expiration on, so that might be a useful > thing > > to do with them. I know Shahar is also looking at using credits for CI > > runners. > > > > > > > On 9 Oct 2026, at 11:28, Shubham Raj <[email protected]> wrote: > > > > > > Hello everyone, > > > > > > As a follow-up to this thread, I've written up the proposal, Here is an > > AIP > > > < > > > https://cwiki.apache.org/confluence/spaces/AIRFLOW/pages/451979463/WIP+AIP-124+Airflow-Powered+Agentic+Workflows+for+Project+Operations > > > > > > , > > > > > > TL;DR: Move our PR triage and review assistance from a "push" model (a > > > maintainer running local agent sessions) to a "pull" model. Airflow > > > continuously discovers work, handles the deterministic parts, and > > surfaces > > > only the decisions that need a human in a shared HITL inbox. This > removes > > > the single point of failure and lets more maintainers share the load. > > > > > > *Key points:* > > > > > > - No new Airflow core subsystem. It reuses existing HITL, common.ai, > > > Assets, deferral, and the official Helm chart (requires Airflow > 3.3+). > > > - Runs in a project-controlled AWS account. Workflows, prompts, > > > policies, and tests live in a new airflow-agentic-workflows repo, so > > anyone > > > can propose changes via PRs. > > > - Bounded event loop, not an open-ended agent. Airflow owns the state > > > machine; the model returns typed results, and every write requiring > > > judgment is gated by a human. > > > - Explicit non-goals: no automatic merging, no unattended approvals > or > > > request-changes, no autonomous workflow approval/rerun. > > > - Rollout starts in shadow mode, comparing results against current > > > maintainer practice before enabling any GitHub writes. > > > > > > The initial scope is PR triage and review. The security issue workflow > > > would follow as a separate track. > > > > > > *Open questions where I'd especially value input:* > > > > > > 1. The initial triage rules (R1-R6). More inputs would help here; are > > > any rules missing or too aggressive? > > > 2. Hosting and operational ownership of the AWS deployment. > > > 3. The attribution model for reviews posted via a central GitHub > > account > > > on behalf of the approving maintainer. > > > 4. Token/model source: Bedrock, llm.apache.org, or something else? > > > 5. What quality gate should we require before leaving shadow mode? > > > > > > Please leave comments here or directly on the wiki page. Once the > > > discussion settles, I'll update the AIP and we can move toward a vote. > > > > > > Thanks, > > > Shubham > > > > > > On Fri, Sep 11, 2026 at 3:02 PM Jarek Potiuk <[email protected]> wrote: > > > > > >> Hi everyone, > > >> > > >> A quick follow-up: something I read today offers a slightly different > > >> perspective on the proposal to use Airflow, which might help clarify > my > > aim: > > >> > > >> Read that*: AWS open-sources Pizza Bot: *email-style inbox for > > background > > >> AI agents - The New Stack: > > >> https://thenewstack.io/aws-pizza-bot-agent-inbox/ > > >> > > >> I want our triage and security process to follow a similar model. > Agents > > >> would work mostly in the background, completing tasks autonomously and > > >> surfacing to maintainers (in an inbox/to-do list fashion) only when > > human > > >> decisions are required. We would have a shared `inbox` style places > > where > > >> maintainers can log in and perform Human-in-the-Loop (HITL) actions. > > >> > > >> Using Airflow to drive this allows us to tailor the experience closely > > to > > >> our needs. It also showcases Airflow for this type of workload and > > >> Dogfooding it offers numerous advantages over using a third-party > > solution > > >> like Pizza Bot. Specifically, we could implement: > > >> > > >> - A shared inbox for the security team to review issues, make triage > > >> decisions, trigger fix PRs, and handle release announcements. > > >> - Per-area inboxes for maintainers to make triage decisions on new > and > > >> outstanding PRs > > >> - Learning what Airflow still needs to support these kinds of > workflows > > >> > > >> This proposal heavily optimizes for minimizing human interaction by > > >> automatically handling straightforward tasks and batching decisions to > > >> surface only high-priority items with full context. It also eliminates > > the > > >> Single Point of Failure (SPOF)—currently myself—by enabling the > broader > > >> community to share the load (but only the "make decisions" part; > Airflow > > >> will handle the execution). > > >> > > >> Quick FAQ: > > >> > > >> *How much Magpie is involved?* > > >> > > >> On the runtime/execution side for Airflow: none. Magpie serves purely > > as a > > >> starting point and interactive testing ground. It already contains > > useful > > >> building blocks, mini-workflows, and integrations that we can learn > > from or > > >> adapt. We can learn from or adapt these integrations and try out new > > things > > >> However, embedding or weaving Magpie into Airflow DAGs is a non-goal. > We > > >> might re-use some code/SKILLs initially, but making this a long-term > > >> dependency is not the goal. Possibly Magpie could eventually implement > > an > > >> "export" feature to Airflow for any project wanting to automate its > > >> triaging/security issue handling with Airflow - following our > blueprint. > > >> However, this is not a goal - it's more of a side effect of using > > similar > > >> features while interactively testing things with Magpie and wanting to > > make > > >> them into a DAG. > > >> > > >> *Should we analyze stats on what works first?* > > >> > > >> Yes. Absolutely. We have, and we will continue analyzing stats while > > >> incrementally moving some workflows to run via Airflow. I have already > > >> gathered and shared these stats from my manual runs with Magpie and we > > can > > >> continue doing so. Look at my past messages where I shared stats, drew > > >> conclusions, and decided on next steps. Magpie provides the tooling to > > run > > >> experiments and evaluate results quickly. Transitioning this to a > > >> continuous, automated Airflow setup removes the single point of > failure > > >> (SPOF) and allows us to iterate and improve these workflows more > > >> collaboratively. Since the decisions we make will have a continuous, > > >> immediate effect once implemented and deployed in our "Airflow > driver," > > we > > >> will hopefully make them more consciously. > > >> > > >> Best regards, > > >> Jarek > > >> > > >> > > >> On Thu, Sep 10, 2026 at 7:02 PM Jarek Potiuk <[email protected]> > wrote: > > >> > > >>> Hello everyone, > > >>> > > >>> Following up on today's Dev call, Shubham will share a proposal > > document > > >>> soon detailing our discussion from yesterday. However, I wanted to > > outline > > >>> the idea here first to clear up a few misconceptions that arose > during > > >>> brief conversations on Slack and with PMC members. If agreed upon, > this > > >>> might eventually become an AIP. > > >>> > > >>> > > >>> *TL;DR: I propose experimenting with dogfooding Airflow > deterministic + > > >>> agentic workflows to handle our triaging and security issue handling* > > >>> > > >>> *Current State* > > >>> We currently use agentic automation via "Magpie" SKILLs ( > > >>> https://magpie.apache.org/) for both areas, with varying success: > > >>> > > >>> - Security issue handling: Working well, processing over 60 issues > > >>> monthly (a 10x increase over six months ago). > > >>> - PR triaging: Partially effective (e.g., filtering bad PRs, with > > >>> potential to automate closing large PRs from inexperienced > > contributors). > > >>> > > >>> *The Problem* > > >>> > > >>> Currently, these workflows require a "push" model where I (in Claude > > CLI) > > >>> initiate actions, confirm proposals, and approve decisions. > Regularly. > > I > > >>> initiate them regularly. Others might also initiate actions, but very > > few > > >>> people have done so. This creates a Single Point of Failure (SPOF). > > IMHO - > > >>> we need to shift to an "automated pull" model to eliminate this SPOF > > and > > >>> allow more community members to act as Human-In-The-Loop (HITL) > > reviewers > > >>> without needing manual initiation. This would also allow more people > to > > >>> contribute to and review those workflows. > > >>> > > >>> *The Proposal* > > >>> > > >>> Instead of relying on local CLI pushes, we can translate these > > learnings > > >>> into Airflow DAGs. Roughly 60%–80% of triage actions are > deterministic > > and > > >>> require no human intervention. For the rest, the automated workflow > > would > > >>> run in the background and trigger "pull" notifications (e.g., Slack > > alerts) > > >>> when bulk human decisions are needed. > > >>> > > >>> *Key Benefits & Features:* > > >>> > > >>> - Shared UI: A cloud-hosted Airflow UI allows multiple maintainers > to > > >>> review and approve decisions. > > >>> - Community-Driven Workflows: Workflows can live in a new Airflow > repo > > >>> (e.g., airflow-agentic-workflows), allowing anyone to modify DAGs via > > pull > > >>> requests. > > >>> - No External Dependencies: While Magpie SKILLs can still be used > for > > >>> local testing or syncs if preferred, using Magpie will not be > required > > for > > >>> PMC/Dev team members. > > >>> - They workflows **might** by somewhat synced with Magpie in either > > way > > >>> (and likely I will start by transplanting parts of the process from > > Magpie > > >>> to Airflow Dags (and implement some features for that in Magpie) . > But > > >>> there will be exactly "0" Magpie components in Airflow Agentic Dags. > > >>> - Dogfooding: Managing Airflow with our latest Helm chart provides > > >>> direct operational feedback for the community—feedback that we > > currently > > >>> miss. > > >>> - We could make this a "use case" demonstrating how Airflow Agentic > > >>> workflows helped solve our problem, which could confirm that Airflow > is > > >>> ideally suited for these kinds of workflows. > > >>> > > >>> *Note on Workflow Architecture (Addressing Vikram's feedback):* > > >>> > > >>> Regarding REACT vs. REPL: Triage and security handling fit the REACT > > >>> pattern (Reason -> Act -> Iterate) due to the continuous stream of > > incoming > > >>> PRs and issues. It's literally -> get issues, evaluate them, make > > actions > > >>> (with or without human) and go back to the beginning. > > >>> REPL relies more on immediate code-execution loops for task > > >>> implementation, whereas our needs align with REACT's iterative > > reasoning > > >>> over time. > > >>> > > >>> Let's discuss further in this thread. Looking forward to your > comments > > >>> and feedback. > > >>> > > >>> Best regards, > > >>> Jarek > > >>> > > >>> > > > > > > --------------------------------------------------------------------- > > To unsubscribe, e-mail: [email protected] > > For additional commands, e-mail: [email protected] > > > > >
