Platform · Benchmarks

The cost of a fix, measured.

We ran five coding agents on real issues from fifteen open-source projects, once with DevOS and once without. Across the counted runs, fix rates were effectively level: 164/206 with DevOS and 163/206 without. The best observed result was a 71% lower cost to fix with DevOS. Every run is published below.

What was run

Seventeen tasks, five models, every run published.

Real issues from fifteen open-source projects in nine languages, worked by five agents across two coding CLIs. Every task, run and exclusion is available here.

17Tasks
15Projects
9Languages
5Models
506Recorded runs
Agent coverage

5 models across 2 coding CLIs.

  • Claude Sonnet 4.6Claude Code
  • Claude Sonnet 5Claude Code
  • Claude Opus 5Claude Code
  • GPT-5.6 SolCodex CLI
  • GPT-5.6 TerraCodex CLI
Run accounting

412 counted. 94 set aside.

Every result uses counted runs. Set-asides remain published with their reason.

412 counted94 set aside3 tries per task, model and side
Results

Each model, with DevOS and without.

Each card is one agent on one kind of task, with DevOS and without. Every run it made is a bar.

Search-heavy fixes

element-web · 772df302JavaScript + element-web · a692fe21JavaScript + EFCorePowerTools · 3418C# + scala-steward · 3594Scala · 3 tries each per side · 12 planned a side

Most of the cost here goes to finding the right change and proving it. Knowing the codebase pays best on these: four of the five agents landed the same fixes for less with DevOS, two of them roughly a third cheaper.

Claude Sonnet 4.6 Claude Code

Stock agent
9 of 12 fixed
$3.98 per fix · $35.79 spent in total
9 fixed · 3 unfixed · 0 set aside
With DevOS
9 of 12 fixed
$2.73 per fix · $24.59 spent in total
9 fixed · 3 unfixed · 0 set aside
same fixes−31% cost per fix

Claude Sonnet 5 Claude Code

Stock agent
7 of 11 fixed
1 of 12 recorded runs set aside
$3.07 per fix · $21.46 spent in total
7 fixed · 4 unfixed · 1 set aside
With DevOS
8 of 11 fixed
1 of 12 recorded runs set aside
$3.13 per fix · $25.07 spent in total
8 fixed · 3 unfixed · 1 set aside
+1 fix+2% cost per fix1 run a side set aside: DevOS was not used

Claude Opus 5 Claude Code

Stock agent
5 of 6 fixed
6 of 12 recorded runs set aside
$3.31 per fix · $16.57 spent in total
5 fixed · 1 unfixed · 6 set aside
With DevOS
6 of 6 fixed
6 of 12 recorded runs set aside
$2.28 per fix · $13.69 spent in total
6 fixed · 0 unfixed · 6 set aside
+1 fix−31% cost per fix6 runs a side set aside: the model already knew these fixes

GPT-5.6 Sol Codex CLI

Stock agent
9 of 11 fixed
1 of 12 recorded runs set aside
$0.77 per fix · $6.94 spent in total
9 fixed · 2 unfixed · 1 set aside
With DevOS
9 of 11 fixed
1 of 12 recorded runs set aside
$0.72 per fix · $6.48 spent in total
9 fixed · 2 unfixed · 1 set aside
same fixes−7% cost per fix1 run a side set aside: DevOS was not used

GPT-5.6 Terra Codex CLI

Stock agent
1 of 4 fixed
8 of 12 recorded runs set aside
$2.38 per fix · $2.38 spent in total
1 fixed · 3 unfixed · 8 set aside
With DevOS
3 of 4 fixed
9 of 13 recorded runs set aside
$0.88 per fix · $2.64 spent in total
3 fixed · 1 unfixed · 9 set aside
+2 fixes−63% cost per fix17 runs set aside: DevOS was not used; grading failed; a run lost its comparison partner

Mid-size fixes

NodeBB · 18c45b44JavaScript + openlibrary · 7bf32385Python + stateless · 644C# + just · 3200Rust + task · 2716Go + PhpSpreadsheet · 4541PHP + pana · 1489Dart + pub-dev · 8884Dart + rector-symfony · 797PHP + chezmoi · 5016Go · 3 tries each per side · 30 planned a side

These tasks take real exploration before the change. Most agents come out level on cost here, and Claude Sonnet 4.6 comes out 28% ahead.

Claude Sonnet 4.6 Claude Code

Stock agent
19 of 26 fixed
5 of 31 recorded runs set aside
$2.68 per fix · $51.01 spent in total
19 fixed · 7 unfixed · 5 set aside
With DevOS
17 of 26 fixed
4 of 30 recorded runs set aside
$1.95 per fix · $33.08 spent in total
17 fixed · 9 unfixed · 4 set aside
−2 fixes−28% cost per fix9 runs set aside: the model already knew these fixes; the run produced no result to grade

Claude Sonnet 5 Claude Code

Stock agent
21 of 27 fixed
3 of 30 recorded runs set aside
$2.45 per fix · $51.38 spent in total
21 fixed · 6 unfixed · 3 set aside
With DevOS
19 of 27 fixed
3 of 30 recorded runs set aside
$2.45 per fix · $46.46 spent in total
19 fixed · 8 unfixed · 3 set aside
−2 fixessame cost per fix3 runs a side set aside: the model already knew these fixes

Claude Opus 5 Claude Code

Stock agent
21 of 27 fixed
3 of 30 recorded runs set aside
$1.42 per fix · $29.77 spent in total
21 fixed · 6 unfixed · 3 set aside
With DevOS
21 of 27 fixed
3 of 30 recorded runs set aside
$1.38 per fix · $29.01 spent in total
21 fixed · 6 unfixed · 3 set aside
same fixes−3% cost per fix3 runs a side set aside: the model already knew these fixes

GPT-5.6 Sol Codex CLI

Stock agent
17 of 27 fixed
3 of 30 recorded runs set aside
$0.93 per fix · $15.78 spent in total
17 fixed · 10 unfixed · 3 set aside
With DevOS
19 of 27 fixed
3 of 30 recorded runs set aside
$0.89 per fix · $16.85 spent in total
19 fixed · 8 unfixed · 3 set aside
+2 fixes−4% cost per fix3 runs a side set aside: the model already knew these fixes

GPT-5.6 Terra Codex CLI

Stock agent
12 of 13 fixed
14 of 27 recorded runs set aside
$0.62 per fix · $7.39 spent in total
12 fixed · 1 unfixed · 14 set aside
With DevOS
12 of 13 fixed
14 of 27 recorded runs set aside
$0.70 per fix · $8.37 spent in total
12 fixed · 1 unfixed · 14 set aside
same fixes+13% cost per fix14 runs a side set aside: DevOS was not used; the model already knew these fixes

Small fixes

openlibrary · d8162c22Python + fluentd · 4655Ruby + rubocop · 13705Ruby · 3 tries each per side · 9 planned a side

A handful of turns, well under a dollar. Jobs this small fit in any agent's head, so the differences here are cents. Most agents still came in level or a little cheaper with DevOS.

Claude Sonnet 4.6 Claude Code

Stock agent
9 of 9 fixed
$0.39 per fix · $3.48 spent in total
9 fixed · 0 unfixed · 0 set aside
With DevOS
9 of 9 fixed
$0.31 per fix · $2.82 spent in total
9 fixed · 0 unfixed · 0 set aside
same fixes−19% cost per fix

Claude Sonnet 5 Claude Code

Stock agent
9 of 9 fixed
$0.40 per fix · $3.60 spent in total
9 fixed · 0 unfixed · 0 set aside
With DevOS
9 of 9 fixed
$0.40 per fix · $3.56 spent in total
9 fixed · 0 unfixed · 0 set aside
same fixes−1% cost per fix

Claude Opus 5 Claude Code

Stock agent
9 of 9 fixed
$0.21 per fix · $1.87 spent in total
9 fixed · 0 unfixed · 0 set aside
With DevOS
9 of 9 fixed
$0.23 per fix · $2.04 spent in total
9 fixed · 0 unfixed · 0 set aside
same fixes+9% cost per fix

GPT-5.6 Sol Codex CLI

Stock agent
9 of 9 fixed
$0.43 per fix · $3.90 spent in total
9 fixed · 0 unfixed · 0 set aside
With DevOS
9 of 9 fixed
$0.37 per fix · $3.37 spent in total
9 fixed · 0 unfixed · 0 set aside
same fixes−14% cost per fix

GPT-5.6 Terra Codex CLI

Stock agent
6 of 6 fixed
3 of 9 recorded runs set aside
$0.25 per fix · $1.52 spent in total
6 fixed · 0 unfixed · 3 set aside
With DevOS
5 of 6 fixed
3 of 9 recorded runs set aside
$0.41 per fix · $2.04 spent in total
5 fixed · 1 unfixed · 3 set aside
−1 fix+61% cost per fix3 runs a side set aside: DevOS was not used
The totals

Every model, all tasks together.

One line per agent, averaged over every task on this page, the deep work and the trivial alike.

ModelFixedwith DevOSCost per fixwith DevOSFixedwithoutCost per fixwithoutCost per fixdifference
Claude Sonnet 4.6Claude Code35/47$1.7337/47$2.44−29%
Claude Sonnet 5Claude Code36/47$2.0937/47$2.07+1%
Claude Opus 5Claude Code36/42$1.2435/42$1.38−10%
GPT-5.6 SolCodex CLI37/47$0.7235/47$0.76−5%
GPT-5.6 TerraCodex CLI20/23$0.6519/23$0.59+10%

Fixed counts are out of each side's completed runs. Green means DevOS cost less per fix.

Beyond the bill

We measure DevOS Memory too.

DevOS Memory connects a capability name to the code, configuration and tests behind it. This benchmark isolates that navigation: the task, model and project memory stay the same; only the feature map changes.

Less spend in all seven languages

Both sides produced the same result in every language. Using DevOS Memory cut spend 34% across the set and up to 82% on Go.

Three runs per side

Each language used the same issue, model and project memory on both sides. The only difference was whether DevOS Memory supplied the feature map.

Results by language

LanguageProjectResult using DevOS Memory
GochezmoiSame fixes · 82% less spend
RustjustSame fixes · 17% less spend
ScalaScala StewardSame fixes · 12% less spend
C#EF Core Power ToolsSame fixes · 11% less spend
RubyRuboCopSame fixes · 6% less spend
Dartpub.devNo fix either side · lower spend
PHPPhpSpreadsheetNo fix either side · lower spend

Compared with the same DevOS setup without the feature map. Three runs per side. Scala used Opus; the other languages used Sonnet.

Testing discipline

What it takes for a run to count.

These tasks are public and their fixes are published. Five rules keep the numbers honest, and every set-aside run is listed with its reason.

No internet during runs

Runs can't reach the internet, so an agent can't look up the fix. Attempts are blocked and logged.

A run that reached the web doesn't count

Some agents can call a web tool that runs on their provider's machines, outside the seal we control. We switch those tools off, then check every run's transcript to confirm the switch held. Any run that still reached the web is set aside, together with its matching run on the other side, whatever it scored.

Answers from memory don't count

If a model reproduces a project's published fix from memory, the run doesn't count, and neither does its matching run on the other side. Claude Opus 5 did this on the element-web tasks in both setups. We checked every one of those runs by hand and set them aside.

A DevOS run has to use DevOS

Runs that skipped DevOS say nothing about it. They are set aside, together with their matching runs on the other side.

Reported straight

Results stay per task and per agent. A one-fix difference across three runs is called a tie whichever side it favours.

Cost basis. Claude Code costs are the provider's own session figures. Codex CLI costs are measured tokens at list prices.

Environments. Both sides run in identical, isolated copies of each project.

Sample size. Every figure carries its run count. A saving gets named when it held across runs, and anything inside the run-to-run spread is reported as even. From 506 runs recorded on 2026-08-06. See every one.

Run the comparison on your own stack.

The same layer that produced these numbers installs over the agents your team already runs, measured against your own baseline from day one.