---
title: Astronomer Data Engineering Benchmark
description: >-
  Data Engineering Benchmark to measure how well coding agents do real
  data-engineering work with Airflow.
date: 2026-09-29T00:00:00.000Z
authors:
  - author: src/content/people/anubhav-sharma.md
  - author: src/content/people/julian-laneve.md
  - author: src/content/people/jignesh-patel.md
tags:
  - Apache Airflow
  - benchmarks
  - AI
categories:
  - engineering-blog
canonical_url: 'https://www.astronomer.io/blog/astronomer-data-engineering-benchmark/'
---
_“The daily sales numbers in the dashboard are wrong. Nothing failed, everything was green. Find out why and fix it.”_

This sounds like a simple data issue. But orchestration changes things. At Astronomer, our customers run data platforms with Airflow at the center. The root cause can now sit several layers behind the visible symptoms. A wrong number may come from bad data or SQL, but it may also come from an incorrect schedule, a sensor reading the wrong interval, or an upstream task that succeeded without producing the right output. The same is true for mechanical sounding tasks like upgrading to a new Airflow version. Behind that request there can be removed imports, incompatible providers, renamed APIs, and having to do these intricate tasks without changing downstream behaviour can turn simple problems complex.

This complexity is the reality of running production data platforms yet data engineering benchmarks stop at the warehouse. [Spider](https://yale-lily.github.io/spider) and [BIRD](https://bird-bench.github.io/) test the SQL. [ADE-bench](https://github.com/dbt-labs/ade-bench) and [data-eng-bench](https://www.snowflake.com/en/blog/engineering/data-eng-bench-data-engineering-agent-benchmark/) test the dbt models. [DataBench](https://aclanthology.org/2024.lrec-main.1179/) tests the analysis. None of them test the thing that runs all of it in production, and on most platforms that thing is Airflow.

We had no way to measure how well an agent handles tasks like these. We needed the measurement because we built Otto, a data engineering agent, and every change to it has to answer one question: did this help, or did it only look like it helped? So we studied what our customers’ platforms look like. How many teams share one, how the conventions differ between them, what the legacy schedulers still running beside Airflow look like, where the bugs come from.

Together with the Carnegie Mellon Database Group, we built a platform modeled on actual production environments and then built [de-bench](https://github.com/astronomer/de-bench) on top of it. 139 tasks, each with a known-good solution that passes and an untouched project that fails. We ran every agent, model and thinking level we could get through it.

We quickly realized that the frontier models have gotten really good. They are able to solve some of the most complex problems. So we rebuilt Otto around a simple principle: don’t get in the model’s way when it already knows what to do. Instead we double down on supplying knowledge and context the model does not have, especially private or up-to-date information. This is why Otto performs better on tasks like upgrades and versioning.

## Cost and accuracy

Our customers primarily use the Anthropic and OpenAI family of models to do data engineering work. We ran each of the frontier models (at the time of writing) across both families and compared the native harness against Otto’s harness along different thinking levels. Every configuration was run on all 139 tasks and plotted below with cost along the X axis, and pass rate along the Y axis.

<DeBenchParetoChart captionPrefix="Figure 1:" />

Otto holds the production-relevant portion of the Pareto frontier, delivering the higher accuracy at the lower cost. Codex costs less further down the chart, primarily with Terra and Luna at the smaller end of the GPT‑5.6 family. That lower cost comes with lower accuracy. Overall, when we compare the same model at the same thinking level, Otto achieves a higher pass rate in 15 of the 17 matched configurations, with gains as large as 11%.

We evaluate harness pass rate using three metrics, each answering a different question:

- **Pass@1** measures single-attempt accuracy: on a single attempt, what fraction of tasks pass?
- **Pass@k** measures coverage: with k trials per task, what fraction of tasks pass at least once?
- **Passᵏ** measures reliability: with k trials per task, what fraction of tasks pass all k? This is a gruelling test as a single miss in k trials counts as failure.

Each task was run three times (k=3). We also ran five trials for Opus 5 to test whether the results for the top performing model held under a stricter reliability standard. More trials expose more of the agents’ variability. They make Passᵏ harder to achieve and give us greater confidence that the observed ranking is not the result of a small sample. The Opus 5 ordering remained unchanged at k=5, with Otto leading Claude Code across all three metrics. Table 1 averages all three pass-rate metrics and cost across thinking levels for each model and harness.

**Table 1: pass rate and cost per trial, averaged over thinking levels, for each model and harness.**

| Model | Harness | Avg Pass@1 | Avg Pass@k | Avg Passᵏ | Avg cost |
|---|---|---|---|---|---|
| Opus 5 (k=5) | Otto | 76.0% | 88.7% | 58.8% | $1.22 |
| Opus 5 (k=3) | Otto | 75.9% | 87.0% | 61.4% | $1.23 |
| Opus 5 (k=5) | Claude Code | 72.8% | 84.4% | 57.6% | $1.95 |
| Opus 5 (k=3) | Claude Code | 73.3% | 83.6% | 60.8% | $2.00 |
| Sol 5.6 (k=3) | Otto | 65.5% | 78.2% | 53.0% | $1.15 |
| Sol 5.6 (k=3) | Codex | 59.1% | 71.9% | 45.1% | $0.84 |
| Sonnet 5 (k=3) | Otto | 57.9% | 72.1% | 43.4% | $1.04 |
| Sonnet 5 (k=3) | Claude Code | 54.0% | 68.2% | 39.4% | $1.52 |

Across the table, Otto improves Passᵏ by 1.2% on Opus 5 (while being cheaper), 7.9 points on Sol 5.6, and 4.0 points on Sonnet 5. This means Otto is more likely to solve a task reliably across repeated runs. Agents are nondeterministic, which makes consistency both important and difficult to measure. Solving a task once does not mean an agent will solve it again. Passᵏ sets a high bar by counting a task only when every trial passes. The score is lower, but the signal is more useful because it reflects the performance users can actually depend on.

Otto also leads Pass@1 by 3.2% on Opus 5, 6.4% on Sol 5.6, and 3.9% on Sonnet 5. The same pattern holds for coverage where Otto has better Pass@k by 4.3% on Opus 5, 6.3% on Sol 5.6, and 3.9% on Sonnet 5. This indicates that, across thinking levels, Otto can successfully handle a wider range of tasks than either comparison harness, even when it cannot yet solve every task consistently.

We also compared cost at the top of the performance frontier for the highest accuracy models: Opus 5, Sol 5.6, Sonnet 5. For every Claude Code and Codex configuration there, we found the least expensive Otto configuration within two percentage points of its pass rate for comparable accuracy (within 2%). Figure 2 shows each pair.

![Cost at the same accuracy](/images/posts/2026/astronomer-data-engineering-benchmark/cost-at-the-same-accuracy.png "Figure 2: harness cost at matched accuracy")

The competing configuration costs between 1.09x and 4.32x the comparable Otto configuration, with one exception where the rival costs 8% less. The difference is largest against Claude Code at the top of the frontier. The least expensive Claude Code configuration within two points of Otto’s Opus 5 configuration at low thinking still costs 2.2x. The gap against Codex is smaller, but Otto can still win on both dimensions: Otto with Sol 5.6 at medium thinking scores 62.3% for $1.10 per trial, while Codex with Sol 5.6 at high thinking scores 61.1% for $1.20.

Otto can call models from either OpenAI or Anthropic, allowing users to choose the strongest model for a particular job instead of being tied to one provider. That flexibility makes it possible to optimize cost and performance together. As open-weight models continue to improve, the same architecture could unlock further price-performance gains by pairing Otto’s domain knowledge with capable models at a lower cost. That is an area we intend to keep watching and testing.

Otto’s performance comes from its domain knowledge and optimization rather than greater speed or scale. It is the skills and tools we built to help it function and perform tasks better. Claude Code carries more general-purpose tooling and context rather than being optimized specifically for data engineering. That difference appears in the cost profile: it carries ~56% more context on every turn than Otto. Codex, by contrast, genuinely earns its lower price on the GPT models and does not carry overpowering extra context.

![Pass@1 range by model tier](/images/posts/2026/astronomer-data-engineering-benchmark/pass-at-1-range-by-model-tier.png "Figure 3: Pass@1 range per model. Each bar spans its best and worst configuration across every harness and thinking level.")

### What this tells us

- Model choice matters most. Figure 3 spans each model’s best and worst configuration, and the bands barely overlap. Stronger models drive the largest accuracy gains. In several comparisons, a stronger model at low thinking also delivers better price-performance than a smaller model pushed harder.
- The harness changes price-performance. The same model can produce different accuracy and cost depending on the harness. Claude Code costs 60% more on Opus 5 and 46% more on Sonnet 5 while delivering lower Pass@1 when compared with Otto on the same models. Codex is cheaper on Sol 5.6 on average, but Otto delivers higher accuracy and can cost less when accuracy is matched.
- More thinking helps, but returns diminish. Scores improve in 11 of the 12 sweeps. Near the top, however, the gains shrink while costs rise. With Otto’s Opus 5, high thinking costs 2.4 times as much as minimal thinking for only three additional accuracy points.

### Upgrades and versioning

Everything above says the model does most of the work and the harness should stay out of the way. For Airflow upgrades and versioning, the harness adds value which general model reasoning cannot compensate for even with the frontier models. So the story changes.

Otto scores higher on all seventeen shared model-thinking tiers, Codex included, by margins nothing else in the benchmark comes near: the smallest is 3.3%, the largest 20.4%. Figure 4 shows every tier.

![Upgrade tasks, Pass@1 by tier](/images/posts/2026/astronomer-data-engineering-benchmark/upgrade-tasks-pass-at-1-by-tier.png "Figure 4: Pass@1 on the 62 upgrade tasks, for the 17 model and thinking tiers both harnesses ran.")

This is the clearest evidence that it is the harness knowledge, rather than additional model reasoning, that drives the improvement. With the model and thinking level held constant, Otto achieves higher upgrade accuracy across all 17 matched configurations. The gap is large enough that Otto’s inexpensive configurations outperform Claude Code’s and Codex’s costly ones outright, not merely at a matched price.

The ability to choose models across providers once again makes the difference especially visible. Claude Code’s most expensive upgrade configuration Opus 5 at xhigh thinking scores 65.6% at $1.58 per trial. Otto with Sol at low thinking scores 68.3% at $0.58, roughly one-third of the cost and at one of the lowest available reasoning settings. Otto’s Opus 5 configuration at minimal thinking reaches 72.0% for $0.45. Even Luna, the least expensive model in the entire sweep, scores 51.6% with Otto for $0.07 per trial, ahead of Codex running Sol at medium thinking, which scores 47.3% for $0.48.

What makes Otto effective on upgrades is not access to a smarter model. It is the ability to look up what actually changed between Airflow versions instead of relying on recall or inference. Airflow upgrades involve many small, release-specific facts: which code must move, where it belongs, what was renamed, and what was removed. A model can reason its way to a plausible fix and still miss the single detail that determines whether the project runs. Otto’s harness curates this knowledge into clear, current rules that the agent can retrieve before making a change. It checks the relevant version-specific facts rather than working from a best guess. In practice, that is the difference between a fix that looks complete and one that actually works and a major reason Otto performs so well on upgrade tasks.

_Reasoning budget can’t substitute for a fact the agent was never trained on, or information that is hard to get._

## Creating the benchmark

Building the benchmark turned out to be a much bigger research and engineering effort than we expected. Before we could write the tasks, we had to understand what a representative data platform actually looks like and what makes a benchmark trustworthy. Then we had to recreate that complexity in a controlled environment with tasks and graders that were difficult, fair, and repeatable. We talk about this more in the following sections.

We built de-bench to measure whether an agent can operate across a real, orchestrated data platform. It has three parts: the world agents work in, the tasks they receive, and the checks that grade their work. All three have to be right. A simplified world, an unclear task, or a loose check produces a score that does not transfer to the platforms people actually run.

Our first version used a smaller data platform, and its scores were compressed at the top. Doubling the platform’s size barely moved them. That taught us that difficulty is not just a property of the task; it emerges from the interaction between the task and the world around it.

Consider “find out why this number is wrong.” In a small environment, that might require reading a handful of DAGs. Across six teams with different conventions, overlapping systems, and incomplete migrations, it becomes a very different problem. This shaped every part of the benchmark: the world, the tasks, and the grading system.

### The world

We wanted the world inside de-bench to resemble the data platforms our customers actually operate, so we began by studying those platforms. We looked at how customers organize their work, how large their platforms become, and where teams and systems overlap. Then we built to that shape: real DAGs, data models, contracts, and tests; multiple teams with conflicting conventions; legacy schedulers running alongside Airflow; and bugs that surface the way they do in production.

The resulting world runs Airflow 3 over a DuckDB warehouse, with dbt handling transformations. Half of our contracted customers run dbt through Airflow, and more than 70% already use Airflow 3. The benchmark contains 114 DAGs and roughly 650 Airflow tasks spread across six team projects. It also includes 182 data models built from 57 raw tables and organized into staging, intermediate, and marts layers. Those models span two dbt projects: a mature project with 625 tests and another with 41 models and no tests at all.

That places the benchmark near the middle of the environments we see in practice. The median contracted Astro customer runs about 70 active DAGs, while the median enterprise customer runs about 225. Our benchmark world is larger than most contracted customers’ estates but smaller than most enterprise estates. Its individual components reflect production environments too. A typical customer deployment contains 11 DAGs, and a typical DAG contains four tasks. Large customers rarely become large by expanding a single project. They add projects, and each new project brings another team with its own conventions. That is why we built de-bench as a multi-team platform rather than one oversized project.

However, scale alone does not make a world realistic. Real platforms also contain governance that an agent must read rather than infer: contracts, a lineage registry, a metric dictionary, and project-specific rules. They contain incomplete migrations and multiple legacy schedulers. We included those too.

We also added AGENTS.md files, reflecting how real repositories provide instructions to the agents working inside them. Customers give their own agents this kind of guidance, yet most benchmarks omit it entirely. We did not, because de-bench is designed to measure what agents encounter day to day.

### The tasks

The benchmark contains 139 tasks representing the work engineers actually do on an orchestrated data platform: building and scheduling models, repairing broken pipelines, tracing data-quality problems, documenting findings, and upgrading Airflow projects. The tasks and their supporting projects are available in the [de-bench repository](https://github.com/astronomer/de-bench), so anyone can inspect exactly what the benchmark asks agents to do.

We studied how data platforms fail in production, including postmortems and the “how did nobody catch this for six months?” incidents. We used those patterns to design faults that behave like real failures rather than isolated coding tasks.

Every task ships with a known-good solution and must prove both directions before it is accepted: the solution passes, and the untouched project fails. Before writing the checks, each grading file also identifies the plausible wrong answers. A check that cannot explain how an agent might fool it is only guessing.

The first 77 tasks are written as tickets you could hand to any data engineer. We categorize them by where the required fix actually lives, not by how the symptom is described. The benchmark contains another 62 tasks focused on Airflow upgrades and version compatibility. These include upgrading projects and providers, replacing removed imports, and ensuring that every DAG still parses with the same task IDs on the target release. Each task comes with its own project pinned to an older Airflow version. Table 2 lists the full breakdown by category.

**Table 2: the benchmark tasks by category, data engineering first, then Airflow upgrade and version awareness.**

| Count | Category | Description |
|---|---|---|
| 51 | Data engineering: Airflow, pipeline bugs | Orchestration logic, sensors, triggers, plain Python and SQL outside dbt |
| 13 | Data engineering: dbt model bugs | A staging or intermediate model computing the wrong number |
| 13 | Data engineering: investigation | The correct answer is a written finding with no code change |
| 45 | Airflow upgrade: upgrades | Real code migration tasks across different projects without changing behaviour. |
| 17 | Airflow upgrade: version awareness | Does the agent know which Airflow and provider versions are compatible, and can it author DAGs with that knowledge? |

**Example data engineering task:**

_“RPT-556 - the channel split is wrong for the week of 25 May. Every run that day and that week finished green. Nothing failed, nothing retried, nothing paged, and the on-call notes close the week as “no failures”. Work out why the 25th came out the way it did, and fix it.”_

The above ticket gives the agent a business-level symptom, not a diagnosis. The underlying cause might be a DAG defect, bad data, or a configuration problem; nothing in the ticket reveals which. The agent must investigate across the platform, isolate the cause, and make the correct change. That is what makes it an orchestration and data-engineering problem rather than a narrowly scoped coding exercise.

**Example Airflow upgrade task:**

_“This project defines a custom HTTP data-service operator used by the DAG. Upgrade the project to Airflow 3.0. Preserve the DataServiceOperator’s public signature and the DAG’s observable behaviour”_

Airflow upgrades bring breaking changes, but the ticket does not enumerate them. The agent has to discover every incompatibility: removed imports, renamed helpers, and APIs that no longer exist. It must resolve all of them while preserving the operator’s public interface and the DAG’s behavior.

Every harness runs the same benchmark: identical tickets, trials, and grading criteria.

### Grading

The harder part of a benchmark is not running the agents. It is knowing whether they actually solved the task. Two things can go wrong, and neither necessarily looks like a bug. Noise can drown out the signal, or a check can be misfitted: underfitted checks accept incorrect work, while overfitted checks reject correct work. In either case, every run still completes and every number still looks clean. We recognized this and designed the grading system to guard against both, then tested it to find where those protections failed.

Agents are non-deterministic, and every additional source of noise compounds that variability. Each task therefore runs multiple trials, and every trial is graded across multiple rounds with a majority decision.

For executable tasks, grading proceeds in stages. The DAG must parse cleanly and complete a run. We test the resulting output, then replay the run to verify idempotency. Expected values are recomputed from the raw tables rather than entered by hand. Most checks are deterministic as a result, and a task counts as solved only if every check that ran passed.

Some investigation tasks cannot be graded through execution because their deliverable is a written finding. Those tasks include a human-authored fact key specifying what a correct answer must establish. A model judge evaluates the response against that key and explains its findings before returning a verdict. We hand-adjudicated every answer and required the judge’s decisions to agree with ours across the full set.

We did not assume that this judge was reliable; we measured its reliability and engineered around its variance. We repeatedly graded the same frozen answers to determine whether the verdicts changed. Each grading run uses multiple rounds. A failing verdict requires agreement from a second model, and a third breaks any split decision.

Grader fit required manual auditing because, from the outside, an unfair grader looks exactly like a difficult task. We therefore examined the answers themselves rather than relying on failure rates. Across every judge-graded task, we refined the rubrics until both directions held: failed answers were genuinely wrong, and passing answers had earned the result.

### Our takeaways

- Frontier models are already very capable. Give one a realistic data-engineering task, and it will often find its way to the right answer.
- The harness determines how efficiently it gets there. With the same models doing the same work, Otto delivers better price-performance than Claude Code.
- Domain knowledge matters most on Airflow-specific work. Reasoning cannot replace a fact the model does not know. With that knowledge built in, Otto led every matched configuration, producing the largest gains in the benchmark.

## What this leaves us with

The benchmark is not just a scoreboard. For Otto, the benchmark is both a regression test and a roadmap. It tells us whether a change actually helped and where the next improvement should go. It has already exposed opportunities to reduce unnecessary context, particularly in how our skills are loaded and used, and we will use those findings to make Otto more efficient without sacrificing accuracy.
