Platform · Benchmarks

The cost of a fix, measured.

We ran five coding agents on real issues from fifteen open-source projects, once with DevOS and once without. Across the counted runs, fix rates were effectively level: 176/210 with DevOS and 178/210 without. Cost per fix was up to 78% lower with DevOS. Every run is published below.

What was run

Seventeen tasks, five models, every run published.

Real issues from fifteen open-source projects in nine languages, worked by five agents across two coding CLIs. Every task, run and exclusion is available here.

17Tasks
15Projects
9Languages
5Models
432Recorded runs
Agent coverage

5 models across 2 coding CLIs.

  • Claude Sonnet 4.6Claude Code
  • Claude Sonnet 5Claude Code
  • Claude Opus 5Claude Code
  • GPT-5.6 SolCodex CLI
  • GPT-5.6 TerraCodex CLI
Run accounting

420 counted. 12 set aside.

Every result uses counted runs. Set-asides remain published with their reason.

420 counted12 set aside3 tries per task, model and side
Key results

What the runs showed.

A handful of numbers that hold up under the full record, each with the story behind it. The complete spread — every model, every task, every run — lives in the details.

Claude Sonnet 4.6
up to −78%
Cost per fix on Homebrewery (JavaScript)

Same three fixes on both sides in every run, 59% cheaper on average — and the observed run range reached 78% lower than the stock agent.

Claude Sonnet 5
−57%
Cost per fix on pana (Dart)

Same three fixes on both sides in every run at 57% lower average spend; the observed run range reached 73% lower than the stock agent.

GPT-5.6 Sol
up to −66%
Turns on fluentd (Ruby)

Same three fixes on both sides in every run, with 33% fewer turns and 26% lower cost on average — and the observed run range reached 66% fewer turns.

GPT-5.6 Sol
up to −62%
Cost per fix on the just command runner (Rust)

Same three fixes on both sides in every run, 36% cheaper on average and a fifth fewer turns — and the observed run range reached 62% lower than the stock agent.

DevOS also buys fixes the stock agents miss. On pana, GPT-5.6 Sol and Claude Opus 5 each fixed all three runs with DevOS where their stock runs fixed two — and Opus 5 used 47% less per run doing it. On zod, Claude Sonnet 4.6 landed an extra fix at 42% lower average run cost and a third fewer turns. Sol with DevOS is also the cheapest agent on the board at $0.67 per fix, and GPT-5.6 Terra landed the same Rector Symfony fixes 12% cheaper with DevOS than without.

Beyond the bill

We measure DevOS Memory too.

DevOS Memory connects a capability name to the code, configuration and tests behind it. This benchmark isolates that navigation: the task, model and project memory stay the same; only the feature map changes.

Less spend in all seven languages

Both sides produced the same result in every language. Using DevOS Memory cut spend 34% across the set and up to 82% on Go.

Three runs per side

Each language used the same issue, model and project memory on both sides. The only difference was whether DevOS Memory supplied the feature map.

Results by language

LanguageProjectResult using DevOS Memory
GochezmoiSame fixes · 82% less spend
RustjustSame fixes · 17% less spend
ScalaScala StewardSame fixes · 12% less spend
C#EF Core Power ToolsSame fixes · 11% less spend
RubyRuboCopSame fixes · 6% less spend
Dartpub.devNo fix either side · lower spend
PHPPhpSpreadsheetNo fix either side · lower spend

Compared with the same DevOS setup without the feature map. Three runs per side. Scala used Opus; the other languages used Sonnet.

Testing discipline

What it takes for a run to count.

These tasks are public and their fixes are published. Five rules keep the numbers honest, and every set-aside run is listed with its reason.

No internet during runs

Runs can't reach the internet, so an agent can't look up the fix. Attempts are blocked and logged.

A run that reached the web doesn't count

Some agents can call a web tool that runs on their provider's machines, outside the seal we control. We switch those tools off, then check every run's transcript to confirm the switch held. Any run that still reached the web is set aside, together with its matching run on the other side, whatever it scored.

Answers from memory don't count

If a model reproduces a project's published fix from memory, the run doesn't count, and neither does its matching run on the other side. Claude Opus 5 did this on the element-web tasks in both setups. We checked every one of those runs by hand and set them aside.

A DevOS run has to use DevOS

Runs that skipped DevOS say nothing about it. They are set aside, together with their matching runs on the other side.

Reported straight

Results stay per task and per agent. A one-fix difference across three runs is called a tie whichever side it favours.

Cost basis. Claude Code costs are the provider's own session figures. Codex CLI costs are measured tokens at list prices.

Environments. Both sides run in identical, isolated copies of each project.

Sample size. Every figure carries its run count. A saving gets named when it held across runs, and anything inside the run-to-run spread is reported as even. From 432 runs recorded on 2026-08-20. See every one.

Run the comparison on your own stack.

The same layer that produced these numbers installs over the agents your team already runs, measured against your own baseline from day one.