Leaderboard v2 — Quality × Cost
Every point is a model through a specific harness — the same model may appear once per CLI, because the scaffold measurably matters (native CLIs outscore the Copilot scaffold by 1.2–2.5 points on every pair tested). Token burn is captured where the agent CLI reports it. The v1 board is grandfathered, frozen as of 2026-07-11. Epoch 3 (current): suite v3 red-start guarantees, median of 3 runs per model/harness with min–max spread whiskers, and fail-to-act ("Acts") as a first-class metric. The epoch-2 board is grandfathered, frozen 2026-07-25.
| # | Model | Harness | Route | Score | Grade | Acts | Runs | Time | Tokens | Cost (run) | $/MTok out |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | claude-opus-4-7 | claude | api_key | 87.7 | A | 100% | 3 (86.7–87.9) | 7.7m | 1,709,928 | $1.97 | $25 |
| 2 | claude-opus-4-8 | claude | api_key | 87.1 | A | 100% | 3 (86.2–87.7) | 7.3m | 815,937 | $1.38 | $25 |
| 3 | claude-haiku-4-5-20251001 | claude | api_key | 86.7 | A | 100% | 3 (86.4–86.8) | 7.0m | 1,797,828 | $0.44 | $5 |
| 4 | claude-sonnet-4.6 | copilot | subscription | 86.6 | A | 100% | 3 (86.3–86.8) | 6.5m | 930,100 | — | flat-fee |
| 5 | claude-opus-5 | copilot | subscription | 86.4 | A | 100% | 3 (85.3–86.4) | 7.7m | 1,830,100 | — | flat-fee |
| 6 | gemini-2.5-flash | gemini | api_key | 85.9 | A | 100% | 3 (75.7–86.9) | 10.5m | — | — | — |
| 7 | claude-opus-4.8 | copilot | subscription | 85.9 | A | 100% | 3 (85.0–86.0) | 6.6m | 1,248,600 | — | flat-fee |
| 8 | claude-sonnet-4-6 | claude | api_key | 85.8 | A | 100% | 3 (85.4–87.4) | 7.7m | 1,292,928 | $0.98 | $15 |
| 9 | claude-fable-5 | claude | api_key | 85.5 | A | 100% | 3 (82.1–85.9) | 7.2m | 852,654 | $2.71 | $50 |
| 10 | claude-opus-5 | claude | api_key | 85.5 | A | 100% | 3 (84.9–86.5) | 11.3m | 1,538,567 | $2.63 | $25 |
| 11 | claude-opus-4.7 | copilot | subscription | 85.4 | A | 100% | 3 (85.2–86.9) | 4.7m | 1,145,800 | — | flat-fee |
| 12 | gpt-5.5 | copilot | subscription | 85.0 | B | 100% | 3 (84.2–85.2) | 9.2m | 1,360,000 | — | flat-fee |
| 13 | gpt-5.4 | copilot | subscription | 84.4 | B | 100% | 3 (84.3–87.7) | 7.4m | 1,164,900 | — | flat-fee |
| 14 | gemini-3.1-pro-preview | gemini | api_key | 84.2 | B | 100% | 3 (75.1–87.0) | 10.6m | — | — | $12 |
| 15 | gpt-5.5 | codex | api_key | 83.7 | B | 100% | 3 (83.5–85.2) | 6.8m | 920,552 | — | $30 |
| 16 | gpt-5.6-luna | codex | api_key | 82.0 | B | 100% | 3 (79.0–82.8) | 3.7m | 443,656 | $0.10† | $1 |
| 17 | gpt-5.6-sol | copilot | subscription | 82.0 | B | 100% | 3 (81.4–86.4) | 5.7m | 836,100 | — | flat-fee |
| 18 | gpt-5.6-terra | codex | api_key | 80.8 | B | 100% | 3 (80.6–81.6) | 4.0m | 439,184 | $1.00† | $12 |
| 19 | gpt-5.6-sol | codex | api_key | 80.7 | B | 100% | 3 (80.2–86.4) | 5.0m | 560,471 | $3.15† | $30 |
| 20 | claude-sonnet-5 | claude | api_key | 80.5 | B | 100% | 3 (80.1–81.7) | 8.6m | 1,744,759 | $1.49 | $10 |
| 21 | gemini-2.0-flash | gemini | api_key | 79.3 | B | 100% | 3 (77.9–86.5) | 12.9m | — | — | — |
| 22 | gemini-2.5-pro | gemini | api_key | 78.8 | B | 100% | 3 (70.1–81.1) | 7.0m | — | — | $10 |
| 23 | gpt-5.3-codex | codex | api_key | 57.1 | C | 75% | 3 (43.7–72.2) | 3.4m | 357,962 | $0.72† | $14 |
Quality = GEval-blended score, refactor-storm suite v3.0.0 (red-start guarantees); median of up to 3 runs per model/harness, spread shown as whiskers. "Acts" = action reliability — the share of behaviors where the agent actually engaged the workspace (fail-to-act is a first-class metric). The same model may appear once per harness — the harness effect is measured, not deduplicated. Cost (run): measured where the CLI reports dollars (claude); † = derived from measured tokens × list per-MTok prices (codex). Gemini CLI reports neither tokens nor dollars, and Copilot bills in AI Credits (no USD conversion) — those rows show no run cost. Cost axis uses list output pricing as of 2026-08-02 until measured token burn covers the board.