Leaderboard v2 — Quality × Cost
Every point is a model through a specific harness — the same model may appear once per CLI, because the scaffold measurably matters (native CLIs outscore the Copilot scaffold by 1.2–2.5 points on every pair tested). Token burn is captured where the agent CLI reports it. The v1 board is grandfathered, frozen as of 2026-07-11. This board is grandfathered (frozen 2026-07-25) — the current epoch-3 board runs suite v3 at K=3.
| # | Model | Harness | Route | Score | Grade | Acts | Time | Tokens | Cost (run) | $/MTok out |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | claude-sonnet-4-6 | claude | api_key | 86.7 | A | — | 8.2m | 1,337,734 | $0.99 | $15 |
| 2 | claude-haiku-4-5-20251001 | claude | api_key | 86.5 | A | — | 6.6m | 1,497,178 | $0.42 | $5 |
| 3 | claude-fable-5 | claude | api_key | 86.5 | A | — | 6.5m | 813,224 | $2.45 | $50 |
| 4 | claude-opus-4.8 | copilot | subscription | 86.4 | A | — | 7.3m | — | — | flat-fee |
| 5 | claude-opus-4.7 | copilot | subscription | 86.4 | A | — | 4.6m | — | — | flat-fee |
| 6 | claude-opus-4-8 | claude | api_key | 86.2 | A | — | 6.8m | 934,269 | $1.39 | $25 |
| 7 | claude-opus-4-6 | claude | api_key | 86.0 | A | — | 7.4m | 981,672 | $1.35 | $25 |
| 8 | gpt-5.4 | copilot | subscription | 85.5 | A | — | 5.1m | — | — | flat-fee |
| 9 | claude-fable-5 | copilot | subscription | 84.8 | B | — | 6.0m | — | — | flat-fee |
| 10 | gpt-5.5 | codex | api_key | 84.7 | B | — | 8.6m | 92,546 | — | $30 |
| 11 | gpt-5.6-sol | copilot | subscription | 84.5 | B | — | 5.3m | — | — | flat-fee |
| 12 | gpt-5.6-sol | codex | api_key | 84.5 | B | — | 3.8m | 69,738 | — | $30 |
| 13 | claude-sonnet-4.6 | copilot | subscription | 84.3 | B | — | 5.5m | — | — | flat-fee |
| 14 | gpt-5.5 | copilot | subscription | 83.1 | B | — | 7.1m | — | — | flat-fee |
| 15 | gpt-5.6-luna | codex | api_key | 82.2 | B | — | 3.4m | 70,521 | — | $1 |
| 16 | claude-sonnet-4-5-20250929 | claude | api_key | 82.2 | B | — | 9.0m | 1,280,869 | $0.98 | $15 |
| 17 | gpt-5.6-terra | codex | api_key | 78.6 | B | — | 3.7m | 70,759 | — | $12 |
| 18 | gpt-5.3-codex | codex | api_key | 66.6 | C | — | 3.6m | 51,216 | — | $14 |
| 19 | o3 | codex | api_key | 64.8 | D | — | 11.7m | — | — | — |
| 20 | gpt-4o-mini | codex | api_key | 36.8 | F | — | 12.4m | — | — | — |
| 21 | gpt-4o | codex | api_key | 32.9 | F | — | 11.5m | — | — | — |
| 22 | gpt-4.1 | codex | api_key | 31.4 | F | — | 4.0m | — | — | — |
Quality = GEval-blended score, refactor-storm suite v2.0.0, single run per model/harness. The same model may appear once per harness — the harness effect is measured, not deduplicated. Cost (run): measured where the CLI reports dollars (claude); † = derived from measured tokens × list per-MTok prices (codex). Gemini CLI reports neither tokens nor dollars, and Copilot bills in AI Credits (no USD conversion) — those rows show no run cost. Cost axis uses list output pricing as of 2026-08-02 until measured token burn covers the board.