Every run, task by task.

The full record behind the benchmark summary: one table per task, grouped by the kind of work, covering all five models on both sides, with each set-aside run listed under its table with the reason.

Results

Every task, every model, both sides.

Solves and cost for each task and model, with DevOS and without. Every cost shows the average of three runs with the cheapest and priciest beside it, and each saving carries the band single runs landed in. Green rows mean DevOS cost less.

Search-heavy fixes

element-web · 772df302JavaScript

element-hq/element-web · SWE-bench Pro
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$1.65$1.33–$1.87$1.65$1.52–$1.770%−23% … +25%
Claude Sonnet 5Claude Code3/32/3$2.80$2.13–$3.45$2.47$2.32–$2.55−13%−48% … +16%
Claude Opus 5Claude Codeno valid runsno valid runsno comparison
GPT-5.6 SolCodex CLI2/33/3$0.58$0.54–$0.62$0.76$0.74–$0.78+23%+17% … +31%
GPT-5.6 TerraCodex CLI1/11/1$0.82$0.82–$0.82$0.88$0.88–$0.88+7%+7% … +7%

Claude Opus 5, with DevOS: 3 of 3 runs set aside: the model reproduced the project’s published fix from memory.

Claude Opus 5, without DevOS: 3 of 3 runs set aside: the model reproduced the project’s published fix from memory.

GPT-5.6 Terra, with DevOS: 2 of 3 runs set aside: the run never actually used DevOS.

GPT-5.6 Terra, without DevOS: 2 of 3 runs set aside: its matching run on the other side was set aside.

element-web · a692fe21JavaScript

element-hq/element-web · SWE-bench Pro
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$1.06$0.79–$1.23$1.10$0.85–$1.40+3%−44% … +44%
Claude Sonnet 5Claude Code2/22/2$0.94$0.87–$1.00$0.85$0.81–$0.89−10%−24% … +3%
Claude Opus 5Claude Codeno valid runsno valid runsno comparison
GPT-5.6 SolCodex CLI3/33/3$0.52$0.47–$0.54$0.67$0.62–$0.76+23%+13% … +37%
GPT-5.6 TerraCodex CLI1/10/1$0.54$0.54–$0.54$0.44$0.44–$0.44−23%−23% … −23%

Claude Sonnet 5, with DevOS: 1 of 3 runs set aside: the run never actually used DevOS.

Claude Sonnet 5, without DevOS: 1 of 3 runs set aside: its matching run on the other side was set aside.

Claude Opus 5, with DevOS: 3 of 3 runs set aside: the model reproduced the project’s published fix from memory.

Claude Opus 5, without DevOS: 3 of 3 runs set aside: the model reproduced the project’s published fix from memory.

GPT-5.6 Terra, with DevOS: 2 of 3 runs set aside: the run never actually used DevOS.

GPT-5.6 Terra, without DevOS: 2 of 3 runs set aside: its matching run on the other side was set aside.

EFCorePowerTools · 3418C#

ErikEJ/EFCorePowerTools · SWE-bench Live
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$1.16$0.98–$1.50$1.51$1.33–$1.87+23%−13% … +47%
Claude Sonnet 5Claude Code3/33/3$2.40$1.17–$3.42$1.37$1.13–$1.73−75%−203% … +33%
Claude Opus 5Claude Code3/33/3$2.08$1.38–$2.59$2.46$2.35–$2.67+16%−10% … +48%
GPT-5.6 SolCodex CLI3/33/3$0.75$0.52–$1.00$0.58$0.48–$0.67−29%−110% … +23%
GPT-5.6 TerraCodex CLIno valid runsno valid runsno comparison

GPT-5.6 Terra, with DevOS: 2 of 3 runs set aside: the run never actually used DevOS.

GPT-5.6 Terra, with DevOS: 1 of 3 runs set aside: the grading step failed, so the run has no result either way.

GPT-5.6 Terra, with DevOS: 1 of 3 runs set aside: no valid run remained on the other side to compare against.

GPT-5.6 Terra, without DevOS: 2 of 3 runs set aside: its matching run on the other side was set aside.

GPT-5.6 Terra, without DevOS: 1 of 3 runs set aside: its matching run on the other side was set aside.

scala-steward · 3594Scala

scala-steward-org/scala-steward · SWE-rebench V2
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code0/30/3$4.33$3.92–$4.90$7.67$5.73–$10.83+44%+15% … +64%
Claude Sonnet 5Claude Code0/30/3$2.54$1.85–$3.40$2.74$2.33–$3.22+7%−46% … +43%
Claude Opus 5Claude Code3/32/3$2.49$1.79–$3.40$3.06$2.09–$4.38+19%−63% … +59%
GPT-5.6 SolCodex CLI1/20/2$0.47$0.45–$0.49$0.45$0.43–$0.48−4%−15% … +5%
GPT-5.6 TerraCodex CLI1/20/2$0.64$0.61–$0.68$0.53$0.50–$0.56−21%−36% … −8%

GPT-5.6 Sol, with DevOS: 1 of 3 runs set aside: the run never actually used DevOS.

GPT-5.6 Sol, without DevOS: 1 of 3 runs set aside: its matching run on the other side was set aside.

GPT-5.6 Terra, with DevOS: 1 of 3 runs set aside: the run never actually used DevOS.

GPT-5.6 Terra, without DevOS: 1 of 3 runs set aside: its matching run on the other side was set aside.

Mid-size fixes

NodeBB · 18c45b44JavaScript

NodeBB/NodeBB · SWE-bench Pro
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$1.93$1.70–$2.35$2.61$1.68–$3.63+26%−39% … +53%
Claude Sonnet 5Claude Code2/33/3$4.08$2.63–$5.13$4.95$4.25–$5.50+18%−21% … +52%
Claude Opus 5Claude Code3/33/3$2.63$2.04–$3.27$2.18$2.03–$2.32−20%−61% … +12%
GPT-5.6 SolCodex CLI0/30/3$0.69$0.53–$0.81$0.75$0.66–$0.83+8%−23% … +35%
GPT-5.6 TerraCodex CLIno valid runsno valid runsno comparison

GPT-5.6 Terra, with DevOS: 3 of 3 runs set aside: the run never actually used DevOS.

GPT-5.6 Terra, without DevOS: 3 of 3 runs set aside: its matching run on the other side was set aside.

openlibrary · 7bf32385Python

internetarchive/openlibrary · SWE-bench Pro
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.58$0.47–$0.67$0.56$0.53–$0.61−3%−26% … +23%
Claude Sonnet 5Claude Code3/33/3$1.01$0.73–$1.23$1.05$0.73–$1.50+4%−67% … +51%
Claude Opus 5Claude Code3/33/3$0.94$0.79–$1.14$0.82$0.67–$0.90−14%−70% … +12%
GPT-5.6 SolCodex CLI3/33/3$0.56$0.50–$0.63$0.49$0.39–$0.64−15%−62% … +22%
GPT-5.6 TerraCodex CLIno valid runsno valid runsno comparison

stateless · 644C#

dotnet-state-machine/stateless · SWE-bench Live
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code1/32/3$0.83$0.57–$1.21$0.95$0.70–$1.19+12%−73% … +52%
Claude Sonnet 5Claude Code2/32/3$1.52$1.45–$1.60$1.26$1.15–$1.35−21%−39% … −7%
Claude Opus 5Claude Code3/33/3$0.90$0.85–$0.97$0.88$0.84–$0.93−2%−16% … +8%
GPT-5.6 SolCodex CLI3/33/3$0.54$0.45–$0.68$0.39$0.35–$0.42−37%−95% … −6%
GPT-5.6 TerraCodex CLIno valid runsno valid runsno comparison

GPT-5.6 Terra, with DevOS: 3 of 3 runs set aside: the run never actually used DevOS.

GPT-5.6 Terra, without DevOS: 3 of 3 runs set aside: its matching run on the other side was set aside.

just · 3200Rust

casey/just · SWE-bench Live
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.78$0.54–$1.06$2.68$0.83–$6.16+71%−27% … +91%
Claude Sonnet 5Claude Code3/33/3$1.40$1.13–$1.58$1.38$1.17–$1.54−2%−35% … +27%
Claude Opus 5Claude Code3/33/3$0.60$0.44–$0.76$0.58$0.36–$0.73−5%−112% … +39%
GPT-5.6 SolCodex CLI3/33/3$0.63$0.56–$0.74$0.78$0.68–$0.93+19%−8% … +39%
GPT-5.6 TerraCodex CLI3/33/3$0.58$0.48–$0.68$0.70$0.61–$0.79+18%−13% … +40%

task · 2716Go

go-task/task · SWE-bench Live
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.87$0.73–$0.99$0.85$0.72–$0.94−2%−38% … +23%
Claude Sonnet 5Claude Code3/33/3$0.60$0.41–$0.88$0.54$0.45–$0.61−11%−94% … +32%
Claude Opus 5Claude Code3/33/3$0.74$0.51–$0.86$0.61$0.47–$0.75−20%−84% … +32%
GPT-5.6 SolCodex CLI3/33/3$0.34$0.32–$0.38$0.32$0.27–$0.41−7%−41% … +22%
GPT-5.6 TerraCodex CLI3/33/3$0.35$0.28–$0.39$0.31$0.18–$0.50−11%−120% … +43%

PhpSpreadsheet · 4541PHP

PHPOffice/PhpSpreadsheet · SWE-rebench V2
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code0/30/3$1.14$1.00–$1.37$1.36$1.01–$1.66+16%−36% … +40%
Claude Sonnet 5Claude Code0/30/3$1.20$0.74–$1.57$1.25$1.08–$1.57+4%−45% … +53%
Claude Opus 5Claude Code0/30/3$1.18$0.89–$1.59$1.09$0.96–$1.29−9%−65% … +31%
GPT-5.6 SolCodex CLI0/30/3$0.69$0.50–$1.01$0.46$0.36–$0.60−52%−181% … +18%
GPT-5.6 TerraCodex CLIno valid runsno valid runsno comparison

GPT-5.6 Terra, with DevOS: 3 of 3 runs set aside: the run never actually used DevOS.

GPT-5.6 Terra, without DevOS: 3 of 3 runs set aside: its matching run on the other side was set aside.

pana · 1489Dart

dart-lang/pana · SWE-rebench V2
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$1.44$1.19–$1.76$1.48$1.08–$1.74+2%−63% … +32%
Claude Sonnet 5Claude Code3/33/3$0.97$0.77–$1.10$2.01$1.08–$3.08+52%−1% … +75%
Claude Opus 5Claude Code3/33/3$0.43$0.38–$0.51$0.91$0.51–$1.21+53%0% … +68%
GPT-5.6 SolCodex CLI3/32/3$0.70$0.51–$0.86$0.65$0.60–$0.69−7%−44% … +25%
GPT-5.6 TerraCodex CLI2/33/3$1.15$1.01–$1.25$0.63$0.59–$0.70−82%−113% … −46%

pub-dev · 8884Dart

dart-lang/pub-dev · SWE-rebench V2
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code0/30/3$2.88$2.16–$4.19$4.03$3.59–$4.69+29%−17% … +54%
Claude Sonnet 5Claude Code0/31/3$4.16$3.17–$5.18$4.07$3.52–$4.80−2%−47% … +34%
Claude Opus 5Claude Code0/30/3$1.90$1.48–$2.43$2.27$1.98–$2.51+16%−23% … +41%
GPT-5.6 SolCodex CLI1/30/3$0.98$0.89–$1.04$1.00$0.80–$1.36+2%−29% … +34%
GPT-5.6 TerraCodex CLI1/10/1$0.81$0.81–$0.81$1.12$1.12–$1.12+27%+27% … +27%

GPT-5.6 Terra, with DevOS: 2 of 3 runs set aside: the run never actually used DevOS.

GPT-5.6 Terra, without DevOS: 2 of 3 runs set aside: its matching run on the other side was set aside.

rector-symfony · 797PHP

rectorphp/rector-symfony · SWE-rebench V2
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Codeno valid runsno valid runsno comparison
Claude Sonnet 5Claude Code3/33/3$0.55$0.37–$0.70$0.62$0.42–$0.82+12%−69% … +54%
Claude Opus 5Claude Code3/33/3$0.36$0.34–$0.39$0.58$0.41–$0.69+39%+4% … +51%
GPT-5.6 SolCodex CLI3/33/3$0.48$0.33–$0.67$0.43$0.27–$0.64−13%−151% … +48%
GPT-5.6 TerraCodex CLI1/11/1$0.39$0.39–$0.39$0.49$0.49–$0.49+20%+20% … +20%

Claude Sonnet 4.6, with DevOS: 2 of 3 runs set aside: the model reproduced the project’s published fix from memory.

Claude Sonnet 4.6, with DevOS: 1 of 3 runs set aside: its matching run on the other side was set aside.

Claude Sonnet 4.6, without DevOS: 1 of 3 runs set aside: the model reproduced the project’s published fix from memory.

Claude Sonnet 4.6, without DevOS: 2 of 3 runs set aside: its matching run on the other side was set aside.

GPT-5.6 Terra, with DevOS: 1 of 3 runs set aside: the run never actually used DevOS.

GPT-5.6 Terra, with DevOS: 1 of 3 runs set aside: the model reproduced the project’s published fix from memory.

GPT-5.6 Terra, without DevOS: 1 of 3 runs set aside: its matching run on the other side was set aside.

GPT-5.6 Terra, without DevOS: 1 of 3 runs set aside: its matching run on the other side was set aside.

chezmoi · 5016Go

twpayne/chezmoi · SWE-bench Live
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code1/22/2$0.87$0.35–$1.39$3.74$1.14–$6.35+77%−22% … +94%
Claude Sonnet 5Claude Codeno valid runsno valid runsno comparison
Claude Opus 5Claude Codeno valid runsno valid runsno comparison
GPT-5.6 SolCodex CLIno valid runsno valid runsno comparison
GPT-5.6 TerraCodex CLI2/22/2$0.48$0.47–$0.49$0.42$0.40–$0.45−12%−22% … −4%

Claude Sonnet 4.6, with DevOS: 1 of 3 runs set aside: its matching run on the other side was set aside.

Claude Sonnet 4.6, without DevOS: 2 of 3 runs set aside: the run produced no result to grade.

Claude Sonnet 5, with DevOS: 2 of 3 runs set aside: the model reproduced the project’s published fix from memory.

Claude Sonnet 5, with DevOS: 1 of 3 runs set aside: its matching run on the other side was set aside.

Claude Sonnet 5, without DevOS: 2 of 3 runs set aside: the model reproduced the project’s published fix from memory.

Claude Sonnet 5, without DevOS: 1 of 3 runs set aside: its matching run on the other side was set aside.

Claude Opus 5, with DevOS: 3 of 3 runs set aside: the model reproduced the project’s published fix from memory.

Claude Opus 5, without DevOS: 3 of 3 runs set aside: the model reproduced the project’s published fix from memory.

GPT-5.6 Sol, with DevOS: 3 of 3 runs set aside: its matching run on the other side was set aside.

GPT-5.6 Sol, without DevOS: 3 of 3 runs set aside: the model reproduced the project’s published fix from memory.

GPT-5.6 Terra, with DevOS: 1 of 3 runs set aside: its matching run on the other side was set aside.

GPT-5.6 Terra, without DevOS: 1 of 3 runs set aside: the model reproduced the project’s published fix from memory.

Small fixes

openlibrary · d8162c22Python

internetarchive/openlibrary · SWE-bench Pro
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.27$0.22–$0.34$0.30$0.24–$0.40+11%−40% … +43%
Claude Sonnet 5Claude Code3/33/3$0.30$0.29–$0.31$0.26$0.21–$0.29−17%−51% … +2%
Claude Opus 5Claude Code3/33/3$0.17$0.15–$0.19$0.19$0.12–$0.25+7%−62% … +41%
GPT-5.6 SolCodex CLI3/33/3$0.39$0.35–$0.43$0.23$0.22–$0.24−73%−100% … −47%
GPT-5.6 TerraCodex CLIno valid runsno valid runsno comparison

GPT-5.6 Terra, with DevOS: 3 of 3 runs set aside: the run never actually used DevOS.

GPT-5.6 Terra, without DevOS: 3 of 3 runs set aside: its matching run on the other side was set aside.

fluentd · 4655Ruby

fluent/fluentd · SWE-bench Multilingual
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.32$0.27–$0.36$0.49$0.32–$0.60+34%−14% … +55%
Claude Sonnet 5Claude Code3/33/3$0.47$0.30–$0.65$0.35$0.32–$0.39−35%−103% … +24%
Claude Opus 5Claude Code3/33/3$0.25$0.23–$0.27$0.29$0.28–$0.29+13%+2% … +21%
GPT-5.6 SolCodex CLI3/33/3$0.44$0.36–$0.50$0.83$0.48–$1.27+47%−6% … +71%
GPT-5.6 TerraCodex CLI3/33/3$0.40$0.37–$0.44$0.27$0.19–$0.36−48%−130% … −3%

rubocop · 13705Ruby

rubocop/rubocop · SWE-bench Multilingual
ModelSolvedwith DevOSSolvedwithoutCostwith DevOS · average and rangeCostwithout · average and rangeDevOS savingaverage · single-run band
Claude Sonnet 4.6Claude Code3/33/3$0.35$0.26–$0.40$0.37$0.29–$0.45+5%−41% … +42%
Claude Sonnet 5Claude Code3/33/3$0.41$0.31–$0.52$0.59$0.35–$0.80+31%−49% … +61%
Claude Opus 5Claude Code3/33/3$0.26$0.20–$0.34$0.15$0.10–$0.20−69%−241% … −3%
GPT-5.6 SolCodex CLI3/33/3$0.29$0.27–$0.31$0.24$0.22–$0.27−20%−42% … +1%
GPT-5.6 TerraCodex CLI2/33/3$0.28$0.26–$0.32$0.24$0.20–$0.27−19%−64% … +5%

“No comparison” means one side has no valid runs to compare; the note under the table says why. A difference of a single solve out of three runs is within run-to-run variation, and we report it as a tie.

Run the comparison on your own stack.

The same layer that produced these numbers installs over the agents your team already runs, measured against your own baseline from day one.