Benchmark menu

Normalized score across public games (0–100)

Showing top 24 of 35 benchmarked models (updates when chart loads)

Scale: relative 0-100

  1. 73.5
    Claude Opus 5
  2. 66.5
    GPT-5.6 Sol
  3. 66.1
    Claude Fable 5
  4. 57.2
    Grok 4.5
  5. 55.8
    Gemini 3.6 Flash
  6. 55.4
    Gemini 3.5 Flash
  7. 55.0
    GPT-5.4
  8. 53.1
    GPT-5.5
  9. 52.8
    Muse Spark 1.2
  10. 47.6
    Kimi K3
  11. 47.4
    Claude Opus 4.8
  12. 46.8
    GPT-5.4 Mini
  13. 46.7
    O3
  14. 43.5
    Qwen3.7 Plus
  15. 43.4
    GPT-5.6 Terra
  16. 43.2
    GLM-5.2
  17. 42.8
    GPT 5
  18. 42.7
    Qwen3.7 Max
  19. 42.4
    Kimi K2.7 Code
  20. 40.9
    Claude Sonnet 5
  21. 40.2
    GPT-5.6 Luna
  22. 39.7
    GPT-5.4 Nano
  23. 39.5
    Grok Build 0.1
  24. 35.6
    GPT-OSS 120B

Model leaderboard

One row per evaluated model; Best is the highest normalized score across that model's evaluated games in this reasoning variant view. Models appear after this reasoning variant has match history.

Reasoning level: Balanced Games: 8
How to read

We ask each AI model to write a player program for every game. DuelLab runs those programs against each other. Higher scores mean stronger, better-supported performance against the current models.

Score is the normalized 0–100 result across public games. Playable averages only games where the model produced a working player program. Codegen shows the share that worked first try / worked after repair or resend / failed. Players counts the player programs behind a row, while Best is its highest single-game score. W / D / L shows rated wins, draws, and losses. The 90% band, Signal, and reasoning spread help show how much evidence supports the score and how much results vary.

Partial n/3 is an official Overall rank with incomplete evidence; n/3 shows how many reasoning settings have scores, so compare it with care. Provisional reasoning variants do not receive an official rank. Under-covered means evidence is sparse. Capped draws reached the shared move limit and still count as draws. Filters can change the visible order, but the Official or Base column keeps the published rank. Std. price is an estimate; hover over a value for its source and pricing details.

Columns

Choose the optional columns shown in this table.

Balanced filter leaderboard for DuelLab Benchmark
Rank Model Score Rank change Last tested Best Players
1Claude Opus 573.5Baseline2026-08-0590.214
2GPT-5.6 Sol66.5Baseline2026-08-2390.18
3Claude Fable 566.1Baseline2026-08-0577.47
4Grok 4.557.2Baseline2026-08-0587.28
5Gemini 3.6 Flash55.8Baseline2026-08-2390.614
6Gemini 3.5 Flash55.4Baseline2026-08-0584.58
7GPT-5.455.0Baseline2026-08-231008
8GPT-5.553.1Baseline2026-08-0573.98
9Muse Spark 1.252.8Baseline2026-08-2379.77
10Kimi K347.6Baseline2026-08-2268.58
11Claude Opus 4.847.4Baseline2026-08-0580.18
12GPT-5.4 Mini46.8Baseline2026-08-0558.38
13O346.7Baseline2026-08-0562.66
14Qwen3.7 Plus43.5Baseline2026-08-0572.56
15GPT-5.6 Terra43.4Baseline2026-08-0560.68
16GLM-5.243.2Baseline2026-07-3166.38
17GPT 542.8Baseline2026-08-0562.97
18Qwen3.7 Max42.7Baseline2026-08-0579.17
19Kimi K2.7 Code42.4Baseline2026-08-0563.16
20Claude Sonnet 540.9Baseline2026-08-2352.48
21GPT-5.6 Luna40.2Baseline2026-08-2361.48
22GPT-5.4 Nano39.7Baseline2026-08-0563.112
23Grok Build 0.139.5Baseline2026-08-0568.08
24GPT-OSS 120B35.6Baseline2026-08-0546.67
25Ox Alpha33.9Baseline2026-08-2342.821
26North Mini Code33.6Baseline2026-08-0550.412
27MiMo-V2.532.3Baseline2026-08-0544.215
28DeepSeek V4 Flash32.1Baseline2026-07-3145.214
29LongCat 2.031.1Baseline2026-08-2345.112
30Nemotron 3 Ultra 550B A55B29.8Baseline2026-08-0561.95
31Gemini 3.1 Flash Lite29.5Baseline2026-08-0545.66
32Laguna S 2.128.3Baseline2026-08-2339.411
33Minimax M328.0Baseline2026-08-0547.87
34Nex N2 Pro27.8Baseline2026-08-0548.65
35Gemini 3.5 Flash Lite27.0Baseline2026-08-2340.414
PClaude Opus 4.5 Provisional52.3Baseline2026-08-2184.38
PMiMo-V2.5-Pro Provisional43.0Baseline2026-08-0569.05
PQwen3.8 Max Provisional40.1Baseline2026-08-2355.38
PGemma 4 31B Provisional38.0Baseline2026-08-0559.54
PStep 3.7 Flash Provisional32.3Baseline2026-08-0551.54
PInkling Small Provisional31.9Baseline2026-08-2347.58
PInkling Provisional24.7Baseline2026-08-2257.46
PLing 3.0 Flash Provisional21.8Baseline2026-08-2243.012

How this is scored

1. Generate players

Each model is asked to create a program that can play every benchmark game.

2. Play matches

The generated programs compete head-to-head, with both players receiving comparable opportunities.

3. Score each game

Results and how certain they are produce a score from 0 to 100 for each game.

4. Combine games

Game scores are combined into the main leaderboard score. A known failure to create a usable player contributes zero.

5. Keep settings clear

Every row uses the reasoning setting selected for this view.

6. Add context

Detailed tables provide uncertainty, match records, and other clues for careful comparison.