Thanks Ash for the quick response and feedback,

Initially, yes (while in shadow mode), I plan to use copy/paste or GH APIs.
But the final goal would be like using a service account creds, which will
automatically post the reviews on the PR when approved, Today when we use
the harness to post the comments, it ends with "Drafted-by: <harness>
(<model>); reviewed by <username> before posting", we will follow the same
thing, it's just the username would be the maintainer's name who approved
the HITL.

Initially, I'm thinking of implementing this for committers+, and we can
decide on scaling later based on guardrails or rules implementation.

Yes, those are good programs, but they are specific to users. I think that
would be really needed for contributors regarding their own PR creation and
related work. I don't think the program provides API keys for developers,
which we need for the setup. Good to know we have Azure credits; this adds
to the viable options.

Thanks,
Shubham

On Fri, Oct 9, 2026 at 4:18 PM Ash Berlin-Taylor <[email protected]> wrote:

> Love the idea of using Airflow to manage the Airflow project.
>
> Some clarifying questions:
>
> > human-approved comments, labels, draft conversion, closure where
> permitted, and review dispositions;
>
> For the initial trail, and in the long term, who or what is making posting
> comments? Initial version is “the human has to copy them etc” (and we could
> make that easier by providing `gh` commands etc) or an “AirflowBot” GH
> user/app etc? (I think I’m fine either way, I’m just unsure)
>
> Who would have access to see the results of these runs? Who would be able
> to  be the Human in the HITL to approve things? Were you thinking everyone
> in Triage, Committers and PMC groups? I think from reading the AIP you have
> it as Committers+?
>
> Another possible triage rule: history of opening PRs either too quickly,
> or to many and then never responding to comments? Might not be a real case
> so might not be worth it.
>
>
> On question 4: Both Anthropic and OpenAI have “OSS” programs where they
> give (rolling, after re-applying) 6 month free things to contributors of
> projects like Airflow. Downside to that is it would be tied to an
> individual user I suspect. We also have a large number of Azure credits
> that have a (now 1year left) expiration on, so that might be a useful thing
> to do with them. I know Shahar is also looking at using credits for CI
> runners.
>
>
> > On 9 Oct 2026, at 11:28, Shubham Raj <[email protected]> wrote:
> >
> > Hello everyone,
> >
> > As a follow-up to this thread, I've written up the proposal, Here is an
> AIP
> > <
> https://cwiki.apache.org/confluence/spaces/AIRFLOW/pages/451979463/WIP+AIP-124+Airflow-Powered+Agentic+Workflows+for+Project+Operations
> >
> > ,
> >
> > TL;DR: Move our PR triage and review assistance from a "push" model (a
> > maintainer running local agent sessions) to a "pull" model. Airflow
> > continuously discovers work, handles the deterministic parts, and
> surfaces
> > only the decisions that need a human in a shared HITL inbox. This removes
> > the single point of failure and lets more maintainers share the load.
> >
> > *Key points:*
> >
> >   - No new Airflow core subsystem. It reuses existing HITL, common.ai,
> >   Assets, deferral, and the official Helm chart (requires Airflow 3.3+).
> >   - Runs in a project-controlled AWS account. Workflows, prompts,
> >   policies, and tests live in a new airflow-agentic-workflows repo, so
> anyone
> >   can propose changes via PRs.
> >   - Bounded event loop, not an open-ended agent. Airflow owns the state
> >   machine; the model returns typed results, and every write requiring
> >   judgment is gated by a human.
> >   - Explicit non-goals: no automatic merging, no unattended approvals or
> >   request-changes, no autonomous workflow approval/rerun.
> >   - Rollout starts in shadow mode, comparing results against current
> >   maintainer practice before enabling any GitHub writes.
> >
> > The initial scope is PR triage and review. The security issue workflow
> > would follow as a separate track.
> >
> > *Open questions where I'd especially value input:*
> >
> >   1. The initial triage rules (R1-R6). More inputs would help here; are
> >   any rules missing or too aggressive?
> >   2. Hosting and operational ownership of the AWS deployment.
> >   3. The attribution model for reviews posted via a central GitHub
> account
> >   on behalf of the approving maintainer.
> >   4. Token/model source: Bedrock, llm.apache.org, or something else?
> >   5. What quality gate should we require before leaving shadow mode?
> >
> > Please leave comments here or directly on the wiki page. Once the
> > discussion settles, I'll update the AIP and we can move toward a vote.
> >
> > Thanks,
> > Shubham
> >
> > On Fri, Sep 11, 2026 at 3:02 PM Jarek Potiuk <[email protected]> wrote:
> >
> >> Hi everyone,
> >>
> >> A quick follow-up: something I read today offers a slightly different
> >> perspective on the proposal to use Airflow, which might help clarify my
> aim:
> >>
> >> Read that*: AWS open-sources Pizza Bot: *email-style inbox for
> background
> >> AI agents - The New Stack:
> >> https://thenewstack.io/aws-pizza-bot-agent-inbox/
> >>
> >> I want our triage and security process to follow a similar model. Agents
> >> would work mostly in the background, completing tasks autonomously and
> >> surfacing to maintainers (in an inbox/to-do list fashion) only when
> human
> >> decisions are required. We would have a shared `inbox` style places
> where
> >> maintainers can log in and perform Human-in-the-Loop (HITL) actions.
> >>
> >> Using Airflow to drive this allows us to tailor the experience closely
> to
> >> our needs. It also showcases Airflow for this type of workload and
> >> Dogfooding it offers numerous advantages over using a third-party
> solution
> >> like Pizza Bot. Specifically, we could implement:
> >>
> >>  - A shared inbox for the security team to review issues, make triage
> >> decisions, trigger fix PRs, and handle release announcements.
> >>  - Per-area inboxes for maintainers to make triage decisions on new and
> >> outstanding PRs
> >>  - Learning what Airflow still needs to support these kinds of workflows
> >>
> >> This proposal heavily optimizes for minimizing human interaction by
> >> automatically handling straightforward tasks and batching decisions to
> >> surface only high-priority items with full context. It also eliminates
> the
> >> Single Point of Failure (SPOF)—currently myself—by enabling the broader
> >> community to share the load (but only the "make decisions" part; Airflow
> >> will handle the execution).
> >>
> >> Quick FAQ:
> >>
> >> *How much Magpie is involved?*
> >>
> >> On the runtime/execution side for Airflow: none. Magpie serves purely
> as a
> >> starting point and interactive testing ground. It already contains
> useful
> >> building blocks, mini-workflows, and integrations that we can learn
> from or
> >> adapt. We can learn from or adapt these integrations and try out new
> things
> >> However, embedding or weaving Magpie into Airflow DAGs is a non-goal. We
> >> might re-use some code/SKILLs initially, but making this a long-term
> >> dependency is not the goal. Possibly Magpie could eventually implement
> an
> >> "export" feature to Airflow for any project wanting to automate its
> >> triaging/security issue handling with Airflow - following our blueprint.
> >> However, this is not a goal - it's more of a side effect of using
> similar
> >> features while interactively testing things with Magpie and wanting to
> make
> >> them into a DAG.
> >>
> >> *Should we analyze stats on what works first?*
> >>
> >> Yes. Absolutely. We have, and we will continue analyzing stats while
> >> incrementally moving some workflows to run via Airflow. I have already
> >> gathered and shared these stats from my manual runs with Magpie and we
> can
> >> continue doing so. Look at my past messages where I shared stats, drew
> >> conclusions, and decided on next steps. Magpie provides the tooling to
> run
> >> experiments and evaluate results quickly. Transitioning this to a
> >> continuous, automated Airflow setup removes the single point of failure
> >> (SPOF) and allows us to iterate and improve these workflows more
> >> collaboratively. Since the decisions we make will have a continuous,
> >> immediate effect once implemented and deployed in our "Airflow driver,"
> we
> >> will hopefully make them more consciously.
> >>
> >> Best regards,
> >> Jarek
> >>
> >>
> >> On Thu, Sep 10, 2026 at 7:02 PM Jarek Potiuk <[email protected]> wrote:
> >>
> >>> Hello everyone,
> >>>
> >>> Following up on today's Dev call, Shubham will share a proposal
> document
> >>> soon detailing our discussion from yesterday. However, I wanted to
> outline
> >>> the idea here first to clear up a few misconceptions that arose during
> >>> brief conversations on Slack and with PMC members. If agreed upon, this
> >>> might eventually become an AIP.
> >>>
> >>>
> >>> *TL;DR: I propose experimenting with dogfooding Airflow deterministic +
> >>> agentic workflows to handle our triaging and security issue handling*
> >>>
> >>> *Current State*
> >>> We currently use agentic automation via "Magpie" SKILLs (
> >>> https://magpie.apache.org/) for both areas, with varying success:
> >>>
> >>>  - Security issue handling: Working well, processing over 60 issues
> >>> monthly (a 10x increase over six months ago).
> >>>  - PR triaging: Partially effective (e.g., filtering bad PRs, with
> >>> potential to automate closing large PRs from inexperienced
> contributors).
> >>>
> >>> *The Problem*
> >>>
> >>> Currently, these workflows require a "push" model where I (in Claude
> CLI)
> >>> initiate actions, confirm proposals, and approve decisions. Regularly.
> I
> >>> initiate them regularly. Others might also initiate actions, but very
> few
> >>> people have done so. This creates a Single Point of Failure (SPOF).
> IMHO -
> >>> we need to shift to an "automated pull" model to eliminate this SPOF
> and
> >>> allow more community members to act as Human-In-The-Loop (HITL)
> reviewers
> >>> without needing manual initiation. This would also allow more people to
> >>> contribute to and review those workflows.
> >>>
> >>> *The Proposal*
> >>>
> >>> Instead of relying on local CLI pushes, we can translate these
> learnings
> >>> into Airflow DAGs. Roughly 60%–80% of triage actions are deterministic
> and
> >>> require no human intervention. For the rest, the automated workflow
> would
> >>> run in the background and trigger "pull" notifications (e.g., Slack
> alerts)
> >>> when bulk human decisions are needed.
> >>>
> >>> *Key Benefits & Features:*
> >>>
> >>>  - Shared UI: A cloud-hosted Airflow UI allows multiple maintainers to
> >>> review and approve decisions.
> >>>  - Community-Driven Workflows: Workflows can live in a new Airflow repo
> >>> (e.g., airflow-agentic-workflows), allowing anyone to modify DAGs via
> pull
> >>> requests.
> >>>  - No External Dependencies: While Magpie SKILLs can still be used for
> >>> local testing or syncs if preferred, using Magpie will not be required
> for
> >>> PMC/Dev team members.
> >>>  - They workflows **might** by somewhat synced with Magpie in either
> way
> >>> (and likely I will start by transplanting parts of the process from
> Magpie
> >>> to Airflow Dags (and implement some features for that in Magpie) . But
> >>> there will be exactly "0" Magpie components in Airflow Agentic Dags.
> >>>  - Dogfooding: Managing Airflow with our latest Helm chart provides
> >>> direct operational feedback for the community—feedback that we
> currently
> >>> miss.
> >>>  - We could make this a "use case" demonstrating how Airflow Agentic
> >>> workflows helped solve our problem, which could confirm that Airflow is
> >>> ideally suited for these kinds of workflows.
> >>>
> >>> *Note on Workflow Architecture (Addressing Vikram's feedback):*
> >>>
> >>> Regarding REACT vs. REPL: Triage and security handling fit the REACT
> >>> pattern (Reason -> Act -> Iterate) due to the continuous stream of
> incoming
> >>> PRs and issues. It's literally -> get issues, evaluate them, make
> actions
> >>> (with or without human) and go back to the beginning.
> >>> REPL relies more on immediate code-execution loops for task
> >>> implementation, whereas our needs align with REACT's iterative
> reasoning
> >>> over time.
> >>>
> >>> Let's discuss further in this thread. Looking forward to your comments
> >>> and feedback.
> >>>
> >>> Best regards,
> >>> Jarek
> >>>
> >>>
>
>
> ---------------------------------------------------------------------
> To unsubscribe, e-mail: [email protected]
> For additional commands, e-mail: [email protected]
>
>

Reply via email to