I thought through the compatibility_date vs $schema differences and I think I 
tend to agree that using $schema would be easier for editors. 
compatibility_date has an edge on user ergonomics, but these days people are 
probably always doing copy-paste or just using an agent to generate the 
skeleton anyway. I’ll switch to $schema and drop compatibility_date.

The parser PR already contains a schema file, and we’ll publish it similar to 
API specs under airflow.apache.org <http://airflow.apache.org/>.

TP


> On 3 Oct 2026, at 16:36, Shahar Epstein <[email protected]> wrote:
> 
> Thanks for sharing it, TP!
> I support the overall direction, specifically the language neutral format.
> For the format to work well with editors and other tools, could we use a
> stable format identifier alongside/instead of compatibility_date, and
> document the schema and semantics in a written spec? It would also help to
> define metadata and typed params in the schema, reserve names for deferred
> features such as assets and task groups, and clarify how run: tasks map to
> code when the YAML is read independently.
> I’d recommend looking at existing open workflow formats for inspiration.
> For example, the Open Workflow Specification
> <https://open-workflow-specification.org/> is more data-centric than
> task-centric, so it may not fit Airflow’s execution model, but its schema
> design could offer useful ideas.
> 
> 
> Shahar
> 
> On Fri, Oct 2, 2026 at 2:04 PM Tzu-ping Chung via dev <
> [email protected]> wrote:
> 
>> Hi all,
>> 
>> With AIP-85, I've been working on a YAML Dag format, and would like some
>> feedback on it. Shortest possible taste first:
>> 
>>    compatibility_date: "2026-10-30"
>>    dag_id: daily_sales
>>    schedule: "@daily"
>>    tasks:
>>      # A classical task using an operator.
>>      # There's no automatic template detection; {{ ds }} not wrapped in
>> $t is a literal.
>>      - id: extract
>>        uses:
>> airflow.providers.amazon.aws.transfers.s3_to_redshift.S3ToRedshiftOperator
>>        with:
>>          s3_bucket: retail-raw
>>          s3_key: {$t: "sales/{{ ds }}.csv"}
>>          schema: public
>>          table: sales
>>      # A taskflow-style task.
>>      # Things inside run: are arguments to the function.
>>      # $x means an XCom input.
>>      - id: notify
>>        run:
>>          channel: "#data"
>>          rows: {$x: extract}
>> 
>> ## Why a new format
>> 
>> DagFactory is the direct inspiration and the feature-parity bar; longer
>> term we'd like this to be the path that supersedes it.
>> 
>> The reason to start fresh rather than extending is that DagFactory is
>> essentially Python transcribed into YAML; import paths, callables named by
>> file+function, a default_args block shaped like a DAG() call. With recent
>> AIP-108 (language SDKs), YAML can provide a more neutral foundation for
>> declaring the dependency structure of tasks implemented in ANY language. We
>> also want to better represent Airflow constructs such as XCom and assets.
>> 
>> == Design guidelines and decisions ==
>> 
>> - Language-neutral, not "Python in YAML". JSON is the data model, YAML
>> just a skin: everything round-trips to JSON.
>> - Literal by default; templating is opt-in with {$t: ...} (inspired by
>> AIP-80). This removes the implicit-Jinja surprises. Also, {$f: ...} is an
>> explicit file template, removing the classic `cat {{ ds }}.sh` ->
>> TemplateNotFound foot-gun. (There's also a $const to mark something as
>> literal explicitly.)
>> - No default_args. Reusable `templates:` composed per task via `extends:`,
>> and the merged task is validated against the real operator. An unaccepted
>> argument is a parse error, not a silent drop.
>> - uses: a fully-qualified operator import path. run: the single code
>> primitive for any language (Python/Go/Java written identically).
>> - compatibility_date (borrowed from Cloudflare Workers) pins the format
>> semantics, so old files keep their meaning as the format evolves.
>> - It ships as an AIP-85 importer in the Task SDK, natively recognized by
>> Airflow (enabled by a configuration).
>> 
>> == Deliberately left out of v1 (all additive, can land later) ==
>> 
>> - Task groups, dynamic task mapping, branching; AIP-104/111/113 (batching,
>> loops, dynamic task groups).
>> - Assets / data-aware scheduling (inlets/outlets, asset expressions,
>> watchers).
>> - Callbacks. They name a callable, so they wait on a declarative
>> callable-reference grammar.
>> - Python run: execution. The format is settled; the @task_handler runtime
>> that runs it is a separate work (Go/Java run: tasks already execute).
>> - Operator shorthand (e.g. amazon.S3ToRedshiftOperator) and cross-file
>> shared config.
>> 
>> I've also created a draft PR that implements the parser. (Not the full
>> AIP-85 importer yet.)
>> https://github.com/apache/airflow/pull/74084
>> 
>> Thanks,
>> TP
>> 
>> 
>> ---------------------------------------------------------------------
>> To unsubscribe, e-mail: [email protected]
>> For additional commands, e-mail: [email protected]
>> 
>> 

Reply via email to