Benchmark menu

Normalized score across public games (0–100)

Showing top 24 of 25 benchmarked models (updates when chart loads)

Scale: relative 0-100

  1. 87.3
    GPT-5.6 Sol
  2. 83.3
    Claude Fable 5
  3. 76.0
    Claude Opus 5
  4. 72.7
    GPT-5.4
  5. 67.5
    GPT-5.5
  6. 66.6
    Claude Opus 4.8
  7. 63.2
    Gemini 3.6 Flash
  8. 62.1
    Grok 4.5
  9. 57.5
    GPT-5.6 Terra
  10. 54.3
    Claude Sonnet 5
  11. 52.8
    GPT-5.6 Luna
  12. 47.3
    Gemini 3.5 Flash
  13. 46.3
    DeepSeek V4 Pro
  14. 46.2
    GPT 5
  15. 44.3
    Ox Alpha
  16. 42.8
    O3
  17. 38.7
    DeepSeek V4 Flash
  18. 35.4
    Gemini 3.5 Flash Lite
  19. 35.0
    Hy3 Preview
  20. 34.8
    GPT-OSS 120B
  21. 31.3
    Nemotron 3 Ultra 550B A55B
  22. 31.0
    MiMo-V2.5
  23. 30.6
    Gemini 3.1 Flash Lite
  24. 29.1
    Step 3.7 Flash

Model leaderboard

One row per evaluated model; Best is the highest normalized score across that model's evaluated games in this reasoning variant view. Models appear after this reasoning variant has match history.

Reasoning level: Intensive Games: 8
How to read

We ask each AI model to write a player program for every game. DuelLab runs those programs against each other. Higher scores mean stronger, better-supported performance against the current models.

Score is the normalized 0–100 result across public games. Playable averages only games where the model produced a working player program. Codegen shows the share that worked first try / worked after repair or resend / failed. Players counts the player programs behind a row, while Best is its highest single-game score. W / D / L shows rated wins, draws, and losses. The 90% band, Signal, and reasoning spread help show how much evidence supports the score and how much results vary.

Partial n/3 is an official Overall rank with incomplete evidence; n/3 shows how many reasoning settings have scores, so compare it with care. Provisional reasoning variants do not receive an official rank. Under-covered means evidence is sparse. Capped draws reached the shared move limit and still count as draws. Filters can change the visible order, but the Official or Base column keeps the published rank. Std. price is an estimate; hover over a value for its source and pricing details.

Columns

Choose the optional columns shown in this table.

Intensive filter leaderboard for DuelLab Benchmark
Rank Model Score Rank change Last tested Best Players
1GPT-5.6 Sol87.3Baseline2026-08-23100 × 28
2Claude Fable 583.3Baseline2026-08-05100 × 28
3Claude Opus 576.0Baseline2026-08-0510013
4GPT-5.472.7Baseline2026-08-2295.17
5GPT-5.567.5Baseline2026-08-0593.87
6Claude Opus 4.866.6Baseline2026-08-0579.57
7Gemini 3.6 Flash63.2Baseline2026-08-2391.612
8Grok 4.562.1Baseline2026-08-2391.38
9GPT-5.6 Terra57.5Baseline2026-08-0576.48
10Claude Sonnet 554.3Baseline2026-08-2383.68
11GPT-5.6 Luna52.8Baseline2026-08-2383.18
12Gemini 3.5 Flash47.3Baseline2026-08-0580.86
13DeepSeek V4 Pro46.3Baseline2026-08-0577.26
14GPT 546.2Baseline2026-08-0567.78
15Ox Alpha44.3Baseline2026-08-2352.123
16O342.8Baseline2026-08-2356.019
17DeepSeek V4 Flash38.7Baseline2026-08-0572.36
18Gemini 3.5 Flash Lite35.4Baseline2026-08-2354.413
19Hy3 Preview35.0Baseline2026-08-0559.37
20GPT-OSS 120B34.8Baseline2026-08-0567.35
21Nemotron 3 Ultra 550B A55B31.3Baseline2026-08-0546.29
22MiMo-V2.531.0Baseline2026-07-3158.65
23Gemini 3.1 Flash Lite30.6Baseline2026-08-0546.45
24Step 3.7 Flash29.1Baseline2026-08-0546.65
25Inkling27.4Baseline2026-08-2360.59
PGLM-5.2 Provisional53.3Baseline2026-08-2253.32
PMuse Spark 1.2 Provisional52.0Baseline2026-08-2153.32
PKimi K3 Provisional48.9Baseline2026-08-2281.05
PClaude Opus 4.5 Provisional43.0Baseline2026-08-2265.28
PQwen3.8 Max Provisional42.4Baseline2026-08-2368.77
PInkling Small Provisional35.2Baseline2026-08-2355.47
PMistral Medium 3.5 Provisional17.8Baseline2026-08-0531.14

How this is scored

1. Generate players

Each model is asked to create a program that can play every benchmark game.

2. Play matches

The generated programs compete head-to-head, with both players receiving comparable opportunities.

3. Score each game

Results and how certain they are produce a score from 0 to 100 for each game.

4. Combine games

Game scores are combined into the main leaderboard score. A known failure to create a usable player contributes zero.

5. Keep settings clear

Every row uses the reasoning setting selected for this view.

6. Add context

Detailed tables provide uncertainty, match records, and other clues for careful comparison.