Thanks Ash for the quick response and feedback, Initially, yes (while in shadow mode), I plan to use copy/paste or GH APIs. But the final goal would be like using a service account creds, which will automatically post the reviews on the PR when approved, Today when we use the harness to post the comments, it ends with "Drafted-by: <harness> (<model>); reviewed by <username> before posting", we will follow the same thing, it's just the username would be the maintainer's name who approved the HITL.
Initially, I'm thinking of implementing this for committers+, and we can decide on scaling later based on guardrails or rules implementation. Yes, those are good programs, but they are specific to users. I think that would be really needed for contributors regarding their own PR creation and related work. I don't think the program provides API keys for developers, which we need for the setup. Good to know we have Azure credits; this adds to the viable options. Thanks, Shubham On Fri, Oct 9, 2026 at 4:18 PM Ash Berlin-Taylor <[email protected]> wrote: > Love the idea of using Airflow to manage the Airflow project. > > Some clarifying questions: > > > human-approved comments, labels, draft conversion, closure where > permitted, and review dispositions; > > For the initial trail, and in the long term, who or what is making posting > comments? Initial version is “the human has to copy them etc” (and we could > make that easier by providing `gh` commands etc) or an “AirflowBot” GH > user/app etc? (I think I’m fine either way, I’m just unsure) > > Who would have access to see the results of these runs? Who would be able > to be the Human in the HITL to approve things? Were you thinking everyone > in Triage, Committers and PMC groups? I think from reading the AIP you have > it as Committers+? > > Another possible triage rule: history of opening PRs either too quickly, > or to many and then never responding to comments? Might not be a real case > so might not be worth it. > > > On question 4: Both Anthropic and OpenAI have “OSS” programs where they > give (rolling, after re-applying) 6 month free things to contributors of > projects like Airflow. Downside to that is it would be tied to an > individual user I suspect. We also have a large number of Azure credits > that have a (now 1year left) expiration on, so that might be a useful thing > to do with them. I know Shahar is also looking at using credits for CI > runners. > > > > On 9 Oct 2026, at 11:28, Shubham Raj <[email protected]> wrote: > > > > Hello everyone, > > > > As a follow-up to this thread, I've written up the proposal, Here is an > AIP > > < > https://cwiki.apache.org/confluence/spaces/AIRFLOW/pages/451979463/WIP+AIP-124+Airflow-Powered+Agentic+Workflows+for+Project+Operations > > > > , > > > > TL;DR: Move our PR triage and review assistance from a "push" model (a > > maintainer running local agent sessions) to a "pull" model. Airflow > > continuously discovers work, handles the deterministic parts, and > surfaces > > only the decisions that need a human in a shared HITL inbox. This removes > > the single point of failure and lets more maintainers share the load. > > > > *Key points:* > > > > - No new Airflow core subsystem. It reuses existing HITL, common.ai, > > Assets, deferral, and the official Helm chart (requires Airflow 3.3+). > > - Runs in a project-controlled AWS account. Workflows, prompts, > > policies, and tests live in a new airflow-agentic-workflows repo, so > anyone > > can propose changes via PRs. > > - Bounded event loop, not an open-ended agent. Airflow owns the state > > machine; the model returns typed results, and every write requiring > > judgment is gated by a human. > > - Explicit non-goals: no automatic merging, no unattended approvals or > > request-changes, no autonomous workflow approval/rerun. > > - Rollout starts in shadow mode, comparing results against current > > maintainer practice before enabling any GitHub writes. > > > > The initial scope is PR triage and review. The security issue workflow > > would follow as a separate track. > > > > *Open questions where I'd especially value input:* > > > > 1. The initial triage rules (R1-R6). More inputs would help here; are > > any rules missing or too aggressive? > > 2. Hosting and operational ownership of the AWS deployment. > > 3. The attribution model for reviews posted via a central GitHub > account > > on behalf of the approving maintainer. > > 4. Token/model source: Bedrock, llm.apache.org, or something else? > > 5. What quality gate should we require before leaving shadow mode? > > > > Please leave comments here or directly on the wiki page. Once the > > discussion settles, I'll update the AIP and we can move toward a vote. > > > > Thanks, > > Shubham > > > > On Fri, Sep 11, 2026 at 3:02 PM Jarek Potiuk <[email protected]> wrote: > > > >> Hi everyone, > >> > >> A quick follow-up: something I read today offers a slightly different > >> perspective on the proposal to use Airflow, which might help clarify my > aim: > >> > >> Read that*: AWS open-sources Pizza Bot: *email-style inbox for > background > >> AI agents - The New Stack: > >> https://thenewstack.io/aws-pizza-bot-agent-inbox/ > >> > >> I want our triage and security process to follow a similar model. Agents > >> would work mostly in the background, completing tasks autonomously and > >> surfacing to maintainers (in an inbox/to-do list fashion) only when > human > >> decisions are required. We would have a shared `inbox` style places > where > >> maintainers can log in and perform Human-in-the-Loop (HITL) actions. > >> > >> Using Airflow to drive this allows us to tailor the experience closely > to > >> our needs. It also showcases Airflow for this type of workload and > >> Dogfooding it offers numerous advantages over using a third-party > solution > >> like Pizza Bot. Specifically, we could implement: > >> > >> - A shared inbox for the security team to review issues, make triage > >> decisions, trigger fix PRs, and handle release announcements. > >> - Per-area inboxes for maintainers to make triage decisions on new and > >> outstanding PRs > >> - Learning what Airflow still needs to support these kinds of workflows > >> > >> This proposal heavily optimizes for minimizing human interaction by > >> automatically handling straightforward tasks and batching decisions to > >> surface only high-priority items with full context. It also eliminates > the > >> Single Point of Failure (SPOF)—currently myself—by enabling the broader > >> community to share the load (but only the "make decisions" part; Airflow > >> will handle the execution). > >> > >> Quick FAQ: > >> > >> *How much Magpie is involved?* > >> > >> On the runtime/execution side for Airflow: none. Magpie serves purely > as a > >> starting point and interactive testing ground. It already contains > useful > >> building blocks, mini-workflows, and integrations that we can learn > from or > >> adapt. We can learn from or adapt these integrations and try out new > things > >> However, embedding or weaving Magpie into Airflow DAGs is a non-goal. We > >> might re-use some code/SKILLs initially, but making this a long-term > >> dependency is not the goal. Possibly Magpie could eventually implement > an > >> "export" feature to Airflow for any project wanting to automate its > >> triaging/security issue handling with Airflow - following our blueprint. > >> However, this is not a goal - it's more of a side effect of using > similar > >> features while interactively testing things with Magpie and wanting to > make > >> them into a DAG. > >> > >> *Should we analyze stats on what works first?* > >> > >> Yes. Absolutely. We have, and we will continue analyzing stats while > >> incrementally moving some workflows to run via Airflow. I have already > >> gathered and shared these stats from my manual runs with Magpie and we > can > >> continue doing so. Look at my past messages where I shared stats, drew > >> conclusions, and decided on next steps. Magpie provides the tooling to > run > >> experiments and evaluate results quickly. Transitioning this to a > >> continuous, automated Airflow setup removes the single point of failure > >> (SPOF) and allows us to iterate and improve these workflows more > >> collaboratively. Since the decisions we make will have a continuous, > >> immediate effect once implemented and deployed in our "Airflow driver," > we > >> will hopefully make them more consciously. > >> > >> Best regards, > >> Jarek > >> > >> > >> On Thu, Sep 10, 2026 at 7:02 PM Jarek Potiuk <[email protected]> wrote: > >> > >>> Hello everyone, > >>> > >>> Following up on today's Dev call, Shubham will share a proposal > document > >>> soon detailing our discussion from yesterday. However, I wanted to > outline > >>> the idea here first to clear up a few misconceptions that arose during > >>> brief conversations on Slack and with PMC members. If agreed upon, this > >>> might eventually become an AIP. > >>> > >>> > >>> *TL;DR: I propose experimenting with dogfooding Airflow deterministic + > >>> agentic workflows to handle our triaging and security issue handling* > >>> > >>> *Current State* > >>> We currently use agentic automation via "Magpie" SKILLs ( > >>> https://magpie.apache.org/) for both areas, with varying success: > >>> > >>> - Security issue handling: Working well, processing over 60 issues > >>> monthly (a 10x increase over six months ago). > >>> - PR triaging: Partially effective (e.g., filtering bad PRs, with > >>> potential to automate closing large PRs from inexperienced > contributors). > >>> > >>> *The Problem* > >>> > >>> Currently, these workflows require a "push" model where I (in Claude > CLI) > >>> initiate actions, confirm proposals, and approve decisions. Regularly. > I > >>> initiate them regularly. Others might also initiate actions, but very > few > >>> people have done so. This creates a Single Point of Failure (SPOF). > IMHO - > >>> we need to shift to an "automated pull" model to eliminate this SPOF > and > >>> allow more community members to act as Human-In-The-Loop (HITL) > reviewers > >>> without needing manual initiation. This would also allow more people to > >>> contribute to and review those workflows. > >>> > >>> *The Proposal* > >>> > >>> Instead of relying on local CLI pushes, we can translate these > learnings > >>> into Airflow DAGs. Roughly 60%–80% of triage actions are deterministic > and > >>> require no human intervention. For the rest, the automated workflow > would > >>> run in the background and trigger "pull" notifications (e.g., Slack > alerts) > >>> when bulk human decisions are needed. > >>> > >>> *Key Benefits & Features:* > >>> > >>> - Shared UI: A cloud-hosted Airflow UI allows multiple maintainers to > >>> review and approve decisions. > >>> - Community-Driven Workflows: Workflows can live in a new Airflow repo > >>> (e.g., airflow-agentic-workflows), allowing anyone to modify DAGs via > pull > >>> requests. > >>> - No External Dependencies: While Magpie SKILLs can still be used for > >>> local testing or syncs if preferred, using Magpie will not be required > for > >>> PMC/Dev team members. > >>> - They workflows **might** by somewhat synced with Magpie in either > way > >>> (and likely I will start by transplanting parts of the process from > Magpie > >>> to Airflow Dags (and implement some features for that in Magpie) . But > >>> there will be exactly "0" Magpie components in Airflow Agentic Dags. > >>> - Dogfooding: Managing Airflow with our latest Helm chart provides > >>> direct operational feedback for the community—feedback that we > currently > >>> miss. > >>> - We could make this a "use case" demonstrating how Airflow Agentic > >>> workflows helped solve our problem, which could confirm that Airflow is > >>> ideally suited for these kinds of workflows. > >>> > >>> *Note on Workflow Architecture (Addressing Vikram's feedback):* > >>> > >>> Regarding REACT vs. REPL: Triage and security handling fit the REACT > >>> pattern (Reason -> Act -> Iterate) due to the continuous stream of > incoming > >>> PRs and issues. It's literally -> get issues, evaluate them, make > actions > >>> (with or without human) and go back to the beginning. > >>> REPL relies more on immediate code-execution loops for task > >>> implementation, whereas our needs align with REACT's iterative > reasoning > >>> over time. > >>> > >>> Let's discuss further in this thread. Looking forward to your comments > >>> and feedback. > >>> > >>> Best regards, > >>> Jarek > >>> > >>> > > > --------------------------------------------------------------------- > To unsubscribe, e-mail: [email protected] > For additional commands, e-mail: [email protected] > >
