Every run, task by task.

The full record behind the benchmark summary: one table per task, grouped by the kind of work, covering all five models on both sides, with each set-aside run listed under its table with the reason.

The aggregate cards

Each model, with DevOS and without.

Each card is one agent on one kind of task, with DevOS and without. Every run it made is a bar. The full task-by-task spread follows below.

Search-heavy fixes

element-web · 772df302JavaScript + element-web · a692fe21JavaScript + EFCorePowerTools · 3418C# + zod · 4377TypeScript · 3 tries each per side · 12 planned a side

Most of the cost here goes to finding the right change and proving it. Knowing the codebase pays best on these: Claude Sonnet 4.6 landed an extra fix at 34% lower cost per fix with DevOS, and Claude Opus 5 landed the same fixes a quarter cheaper.

Claude Sonnet 4.6 Claude Code

Stock agent
7 of 9 fixed
$2.51 per fix · $17.54 spent in total
7 fixed · 2 unfixed · 0 set aside
With DevOS
8 of 9 fixed
3 of 12 recorded runs set aside
$1.65 per fix · $13.18 spent in total
8 fixed · 1 unfixed · 3 set aside
+1 fix−34% cost per fix3 runs set aside: a run lost its comparison partner

Claude Sonnet 5 Claude Code

Stock agent
11 of 12 fixed
$1.84 per fix · $20.24 spent in total
11 fixed · 1 unfixed · 0 set aside
With DevOS
11 of 12 fixed
$2.18 per fix · $23.98 spent in total
11 fixed · 1 unfixed · 0 set aside
same fixes+18% cost per fix

Claude Opus 5 Claude Code

Stock agent
6 of 6 fixed
$1.85 per fix · $11.11 spent in total
6 fixed · 0 unfixed · 0 set aside
With DevOS
6 of 6 fixed
$1.36 per fix · $8.18 spent in total
6 fixed · 0 unfixed · 0 set aside
same fixes−26% cost per fix

GPT-5.6 Sol Codex CLI

Stock agent
9 of 12 fixed
$0.91 per fix · $8.17 spent in total
9 fixed · 3 unfixed · 0 set aside
With DevOS
9 of 12 fixed
$1.16 per fix · $10.46 spent in total
9 fixed · 3 unfixed · 0 set aside
same fixes+28% cost per fix

GPT-5.6 Terra Codex CLI

Stock agent
5 of 9 fixed
$1.33 per fix · $6.67 spent in total
5 fixed · 4 unfixed · 0 set aside
With DevOS
3 of 9 fixed
$2.07 per fix · $6.22 spent in total
3 fixed · 6 unfixed · 0 set aside
−2 fixes+55% cost per fix

Mid-size fixes

openlibrary · 7bf32385Python + just · 3200Rust + task · 2716Go + PhpSpreadsheet · 4541PHP + pana · 1489Dart + rector-symfony · 797PHP + chess.js · 546TypeScript + homebrewery · 4014JavaScript · 3 tries each per side · 24 planned a side

These tasks take real exploration before the change. Every Claude model came out ahead on cost per fix here, up to 14% for Claude Opus 5.

Claude Sonnet 4.6 Claude Code

Stock agent
12 of 12 fixed
$0.92 per fix · $11.07 spent in total
12 fixed · 0 unfixed · 0 set aside
With DevOS
12 of 12 fixed
6 of 18 recorded runs set aside
$0.84 per fix · $10.12 spent in total
12 fixed · 0 unfixed · 6 set aside
same fixes−9% cost per fix6 runs set aside: a run lost its comparison partner

Claude Sonnet 5 Claude Code

Stock agent
18 of 21 fixed
$1.39 per fix · $25.05 spent in total
18 fixed · 3 unfixed · 0 set aside
With DevOS
16 of 21 fixed
$1.30 per fix · $20.78 spent in total
16 fixed · 5 unfixed · 0 set aside
−2 fixes−7% cost per fix

Claude Opus 5 Claude Code

Stock agent
15 of 21 fixed
$1.16 per fix · $17.46 spent in total
15 fixed · 6 unfixed · 0 set aside
With DevOS
15 of 21 fixed
$1.00 per fix · $14.94 spent in total
15 fixed · 6 unfixed · 0 set aside
same fixes−14% cost per fix

GPT-5.6 Sol Codex CLI

Stock agent
14 of 21 fixed
$0.71 per fix · $9.97 spent in total
14 fixed · 7 unfixed · 0 set aside
With DevOS
15 of 21 fixed
$0.73 per fix · $11.01 spent in total
15 fixed · 6 unfixed · 0 set aside
+1 fix+3% cost per fix

GPT-5.6 Terra Codex CLI

Stock agent
13 of 18 fixed
$0.79 per fix · $10.30 spent in total
13 fixed · 5 unfixed · 0 set aside
With DevOS
13 of 18 fixed
$1.07 per fix · $13.95 spent in total
13 fixed · 5 unfixed · 0 set aside
same fixes+35% cost per fix

Small fixes

openlibrary · d8162c22Python + fluentd · 4655Ruby + rubocop · 13705Ruby + prometheus · 13845Go + super_editor · 2787Dart · 3 tries each per side · 15 planned a side

A handful of turns, well under a dollar. Jobs this small fit in any agent's head, so the differences here are cents either way.

Claude Sonnet 4.6 Claude Code

Stock agent
12 of 12 fixed
$0.37 per fix · $4.46 spent in total
12 fixed · 0 unfixed · 0 set aside
With DevOS
12 of 12 fixed
3 of 15 recorded runs set aside
$0.31 per fix · $3.71 spent in total
12 fixed · 0 unfixed · 3 set aside
same fixes−17% cost per fix3 runs set aside: a run lost its comparison partner

Claude Sonnet 5 Claude Code

Stock agent
15 of 15 fixed
$0.41 per fix · $6.19 spent in total
15 fixed · 0 unfixed · 0 set aside
With DevOS
15 of 15 fixed
$0.42 per fix · $6.32 spent in total
15 fixed · 0 unfixed · 0 set aside
same fixes+2% cost per fix

Claude Opus 5 Claude Code

Stock agent
15 of 15 fixed
$0.26 per fix · $3.85 spent in total
15 fixed · 0 unfixed · 0 set aside
With DevOS
15 of 15 fixed
$0.29 per fix · $4.35 spent in total
15 fixed · 0 unfixed · 0 set aside
same fixes+13% cost per fix

GPT-5.6 Sol Codex CLI

Stock agent
15 of 15 fixed
$0.29 per fix · $4.40 spent in total
15 fixed · 0 unfixed · 0 set aside
With DevOS
15 of 15 fixed
$0.32 per fix · $4.82 spent in total
15 fixed · 0 unfixed · 0 set aside
same fixes+10% cost per fix

GPT-5.6 Terra Codex CLI

Stock agent
11 of 12 fixed
$0.30 per fix · $3.27 spent in total
11 fixed · 1 unfixed · 0 set aside
With DevOS
11 of 12 fixed
$0.55 per fix · $6.07 spent in total
11 fixed · 1 unfixed · 0 set aside
same fixes+86% cost per fix
Results

Every task, every model, both sides.

Solves and cost for each task and model, with DevOS and without. Every cost shows the average of three runs with the cheapest and priciest beside it, and each saving carries the band single runs landed in. Green rows mean DevOS cost less.

Search-heavy fixes

element-web · 772df302JavaScript

element-hq/element-web · SWE-bench Pro
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$1.76$1.42–$2.30$1.71$1.51–$1.88−3%−52% … +25%
Claude Sonnet 5Claude Code2/33/3$2.40$2.20–$2.72$2.22$1.88–$2.63−9%−44% … +16%
Claude Opus 5Claude Codeno valid runs†no valid runs†——no comparison
GPT-5.6 SolCodex CLI3/33/3$1.03$0.70–$1.45$0.65$0.57–$0.73−57%−156% … +3%
GPT-5.6 TerraCodex CLI2/33/3$0.91$0.76–$1.06$0.94$0.86–$1.05+3%−24% … +27%

element-web · a692fe21JavaScript

element-hq/element-web · SWE-bench Pro
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.91$0.88–$0.96$1.20$1.18–$1.21+24%+19% … +27%
Claude Sonnet 5Claude Code3/32/3$0.80$0.72–$0.91$0.80$0.77–$0.840%−18% … +14%
Claude Opus 5Claude Codeno valid runs†no valid runs†——no comparison
GPT-5.6 SolCodex CLI3/33/3$0.73$0.41–$1.09$0.75$0.61–$0.95+2%−78% … +57%
GPT-5.6 TerraCodex CLI1/32/3$0.54$0.47–$0.60$0.52$0.48–$0.60−3%−26% … +21%

EFCorePowerTools · 3418C#

ErikEJ/EFCorePowerTools · SWE-bench Live
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Codeno valid runs†no valid runs†——no comparison
Claude Sonnet 5Claude Code3/33/3$2.34$1.05–$3.13$1.51$1.08–$1.82−55%−189% … +42%
Claude Opus 5Claude Code3/33/3$1.90$1.17–$2.37$2.41$2.02–$2.69+21%−18% … +56%
GPT-5.6 SolCodex CLI3/33/3$0.72$0.61–$0.81$0.58$0.39–$0.77−24%−111% … +21%
GPT-5.6 TerraCodex CLIno valid runs†no valid runs†——no comparison

† Claude Sonnet 4.6, with DevOS: 3 of 3 runs set aside: no valid run remained on the other side to compare against.

zod · 4377TypeScript

colinhacks/zod · SWE-rebench V2
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code2/31/3$1.72$1.36–$2.26$2.94$2.48–$3.77+42%+9% … +64%
Claude Sonnet 5Claude Code3/33/3$2.45$2.26–$2.69$2.23$1.72–$2.55−10%−56% … +12%
Claude Opus 5Claude Code3/33/3$0.82$0.79–$0.88$1.29$0.76–$1.74+36%−15% … +55%
GPT-5.6 SolCodex CLI0/30/3$1.01$0.62–$1.21$0.74$0.42–$1.11−36%−189% … +44%
GPT-5.6 TerraCodex CLI0/30/3$0.63$0.56–$0.73$0.76$0.63–$0.84+18%−16% … +33%

Mid-size fixes

openlibrary · 7bf32385Python

internetarchive/openlibrary · SWE-bench Pro
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Codeno valid runs†no valid runs†——no comparison
Claude Sonnet 5Claude Codeno valid runs†no valid runs†——no comparison
Claude Opus 5Claude Codeno valid runs†no valid runs†——no comparison
GPT-5.6 SolCodex CLIno valid runs†no valid runs†——no comparison
GPT-5.6 TerraCodex CLIno valid runs†no valid runs†——no comparison

just · 3200Rust

casey/just · SWE-bench Live
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Codeno valid runs†no valid runs†——no comparison
Claude Sonnet 5Claude Code3/33/3$1.09$0.69–$1.44$1.08$0.97–$1.20−2%−48% … +43%
Claude Opus 5Claude Code3/33/3$0.53$0.40–$0.64$0.66$0.51–$0.93+19%−26% … +57%
GPT-5.6 SolCodex CLI3/33/3$0.43$0.31–$0.52$0.68$0.59–$0.81+36%+12% … +62%
GPT-5.6 TerraCodex CLI3/33/3$0.67$0.58–$0.71$0.53$0.31–$0.67−27%−127% … +13%

† Claude Sonnet 4.6, with DevOS: 3 of 3 runs set aside: no valid run remained on the other side to compare against.

task · 2716Go

go-task/task · SWE-bench Live
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.66$0.47–$0.93$0.69$0.40–$1.02+5%−134% … +54%
Claude Sonnet 5Claude Code3/33/3$0.68$0.56–$0.80$0.50$0.46–$0.53−37%−76% … −5%
Claude Opus 5Claude Code3/33/3$0.46$0.35–$0.59$0.67$0.43–$0.84+31%−38% … +58%
GPT-5.6 SolCodex CLI3/33/3$0.42$0.32–$0.48$0.24$0.17–$0.31−78%−188% … −3%
GPT-5.6 TerraCodex CLI3/33/3$0.48$0.43–$0.56$0.29$0.25–$0.33−62%−121% … −29%

PhpSpreadsheet · 4541PHP

PHPOffice/PhpSpreadsheet · SWE-rebench V2
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Codeno valid runs†no valid runs†——no comparison
Claude Sonnet 5Claude Code0/30/3$1.76$1.57–$2.03$1.41$1.19–$1.66−25%−70% … +6%
Claude Opus 5Claude Code0/30/3$1.41$0.94–$2.17$1.13$1.03–$1.19−25%−110% … +21%
GPT-5.6 SolCodex CLI0/30/3$0.39$0.34–$0.41$0.47$0.34–$0.53+18%−21% … +36%
GPT-5.6 TerraCodex CLIno valid runs†no valid runs†——no comparison

† Claude Sonnet 4.6, with DevOS: 3 of 3 runs set aside: no valid run remained on the other side to compare against.

pana · 1489Dart

dart-lang/pana · SWE-rebench V2
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$1.33$1.28–$1.39$0.98$0.78–$1.28−35%−78% … 0%
Claude Sonnet 5Claude Code3/33/3$1.06$1.05–$1.08$2.49$1.31–$3.87+57%+18% … +73%
Claude Opus 5Claude Code3/32/3$0.58$0.50–$0.69$1.09$0.91–$1.37+47%+24% … +64%
GPT-5.6 SolCodex CLI3/32/3$0.82$0.51–$1.01$0.59$0.45–$0.68−37%−126% … +24%
GPT-5.6 TerraCodex CLI1/31/3$1.48$1.14–$1.94$0.70$0.45–$1.03−110%−335% … −11%

rector-symfony · 797PHP

rectorphp/rector-symfony · SWE-rebench V2
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Codeno valid runs†no valid runs†——no comparison
Claude Sonnet 5Claude Code3/33/3$0.52$0.48–$0.55$0.55$0.52–$0.58+5%−5% … +17%
Claude Opus 5Claude Code3/33/3$0.45$0.37–$0.58$0.63$0.60–$0.66+29%+4% … +43%
GPT-5.6 SolCodex CLI3/33/3$0.43$0.36–$0.54$0.35$0.24–$0.41−24%−123% … +12%
GPT-5.6 TerraCodex CLI3/33/3$0.30$0.27–$0.35$0.35$0.26–$0.52+12%−38% … +47%

chess.js · 546TypeScript

jhlywa/chess.js · SWE-rebench V2
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.93$0.85–$1.06$0.88$0.82–$1.01−5%−29% … +16%
Claude Sonnet 5Claude Code3/33/3$1.08$0.79–$1.34$1.64$1.20–$2.28+34%−11% … +65%
Claude Opus 5Claude Code3/33/3$1.10$0.95–$1.19$1.00$0.49–$1.57−10%−142% … +40%
GPT-5.6 SolCodex CLI3/33/3$0.60$0.51–$0.77$0.58$0.41–$0.80−5%−89% … +36%
GPT-5.6 TerraCodex CLI3/33/3$1.06$0.83–$1.33$0.91$0.72–$1.21−17%−85% … +31%

homebrewery · 4014JavaScript

naturalcrit/homebrewery · SWE-rebench V2
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.46$0.36–$0.59$1.13$0.53–$1.65+59%−10% … +78%
Claude Sonnet 5Claude Code1/33/3$0.74$0.58–$0.83$0.70$0.65–$0.73−5%−29% … +21%
Claude Opus 5Claude Code0/31/3$0.45$0.40–$0.55$0.64$0.35–$0.81+31%−55% … +51%
GPT-5.6 SolCodex CLI0/30/3$0.58$0.49–$0.64$0.42$0.31–$0.57−38%−107% … +13%
GPT-5.6 TerraCodex CLI0/30/3$0.66$0.56–$0.76$0.65$0.49–$0.90−1%−56% … +38%

Small fixes

openlibrary · d8162c22Python

internetarchive/openlibrary · SWE-bench Pro
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.23$0.18–$0.29$0.27$0.24–$0.28+14%−23% … +36%
Claude Sonnet 5Claude Code3/33/3$0.41$0.32–$0.56$0.31$0.24–$0.37−30%−128% … +15%
Claude Opus 5Claude Code3/33/3$0.16$0.14–$0.20$0.18$0.13–$0.21+8%−52% … +33%
GPT-5.6 SolCodex CLI3/33/3$0.34$0.32–$0.38$0.24$0.20–$0.26−45%−83% … −27%
GPT-5.6 TerraCodex CLIno valid runs†no valid runs†——no comparison

fluentd · 4655Ruby

fluent/fluentd · SWE-bench Multilingual
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.32$0.32–$0.33$0.34$0.23–$0.46+7%−44% … +30%
Claude Sonnet 5Claude Code3/33/3$0.37$0.26–$0.43$0.24$0.23–$0.25−52%−86% … −3%
Claude Opus 5Claude Code3/33/3$0.33$0.30–$0.36$0.27$0.23–$0.34−21%−58% … +13%
GPT-5.6 SolCodex CLI3/33/3$0.35$0.33–$0.36$0.48$0.25–$0.84+26%−45% … +60%
GPT-5.6 TerraCodex CLI3/33/3$0.89$0.53–$1.36$0.21$0.19–$0.24−319%−607% … −125%

rubocop · 13705Ruby

rubocop/rubocop · SWE-bench Multilingual
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Codeno valid runs†no valid runs†——no comparison
Claude Sonnet 5Claude Code3/33/3$0.45$0.44–$0.46$0.49$0.46–$0.54+9%+1% … +19%
Claude Opus 5Claude Code3/33/3$0.24$0.15–$0.31$0.11$0.09–$0.13−111%−249% … −16%
GPT-5.6 SolCodex CLI3/33/3$0.29$0.24–$0.38$0.23$0.17–$0.31−26%−122% … +25%
GPT-5.6 TerraCodex CLI2/32/3$0.35$0.33–$0.37$0.21$0.20–$0.22−67%−82% … −51%

† Claude Sonnet 4.6, with DevOS: 3 of 3 runs set aside: no valid run remained on the other side to compare against.

prometheus · 13845Go

prometheus/prometheus · SWE-bench Multilingual
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.32$0.20–$0.39$0.20$0.12–$0.28−62%−211% … +28%
Claude Sonnet 5Claude Code3/33/3$0.40$0.36–$0.44$0.46$0.29–$0.60+14%−51% … +39%
Claude Opus 5Claude Code3/33/3$0.33$0.28–$0.38$0.22$0.17–$0.26−51%−127% … −7%
GPT-5.6 SolCodex CLI3/33/3$0.31$0.30–$0.32$0.20$0.17–$0.25−53%−90% … −21%
GPT-5.6 TerraCodex CLI3/33/3$0.31$0.21–$0.43$0.19$0.13–$0.29−68%−227% … +27%

super_editor · 2787Dart

superlistapp/super_editor · SWE-rebench V2
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.37$0.34–$0.38$0.68$0.40–$1.22+46%+6% … +72%
Claude Sonnet 5Claude Code3/33/3$0.48$0.36–$0.62$0.55$0.38–$0.73+12%−64% … +51%
Claude Opus 5Claude Code3/33/3$0.38$0.38–$0.38$0.50$0.42–$0.58+23%+8% … +35%
GPT-5.6 SolCodex CLI3/33/3$0.32$0.27–$0.35$0.32$0.25–$0.43+3%−42% … +37%
GPT-5.6 TerraCodex CLI3/33/3$0.48$0.42–$0.53$0.48$0.27–$0.65+2%−92% … +36%

“No comparison” means one side has no valid runs to compare; the note under the table says why. A difference of a single solve out of three runs is within run-to-run variation, and we report it as a tie.

Run the comparison on your own stack.

The same layer that produced these numbers installs over the agents your team already runs, measured against your own baseline from day one.