Benchmark menu

Normalized score across public games (0–100)

Showing top 24 of 33 benchmarked models (updates when chart loads)

Scale: relative 0-100

  1. 68.5
    Claude Opus 5
  2. 55.8
    GPT-5.6 Sol
  3. 53.4
    Kimi K3
  4. 49.0
    Claude Opus 4.5
  5. 47.5
    Claude Opus 4.8
  6. 46.4
    Grok 4.5
  7. 43.5
    GPT-5.4
  8. 42.8
    GLM-5.2
  9. 42.1
    Qwen3.7 Max
  10. 41.9
    GPT-5.6 Terra
  11. 41.5
    GPT 4.1
  12. 41.2
    Claude Sonnet 5
  13. 40.8
    DeepSeek V4 Pro
  14. 40.6
    Gemini 3.5 Flash Lite
  15. 40.2
    Muse Spark 1.2
  16. 39.2
    GPT-5.5
  17. 38.0
    Laguna S 2.1
  18. 37.8
    Gemini 3.6 Flash
  19. 37.6
    Minimax M3
  20. 37.2
    Hy3 Preview
  21. 36.1
    GPT-5.4 Nano
  22. 34.9
    MiMo-V2.5
  23. 34.6
    Nemotron 3 Ultra 550B A55B
  24. 34.0
    Inkling

Model leaderboard

One row per evaluated model; Best is the highest normalized score across that model's evaluated games in this reasoning variant view. Models appear after this reasoning variant has match history.

Reasoning level: Baseline Games: 8
How to read

We ask each AI model to write a player program for every game. DuelLab runs those programs against each other. Higher scores mean stronger, better-supported performance against the current models.

Score is the normalized 0–100 result across public games. Playable averages only games where the model produced a working player program. Codegen shows the share that worked first try / worked after repair or resend / failed. Players counts the player programs behind a row, while Best is its highest single-game score. W / D / L shows rated wins, draws, and losses. The 90% band, Signal, and reasoning spread help show how much evidence supports the score and how much results vary.

Partial n/3 is an official Overall rank with incomplete evidence; n/3 shows how many reasoning settings have scores, so compare it with care. Provisional reasoning variants do not receive an official rank. Under-covered means evidence is sparse. Capped draws reached the shared move limit and still count as draws. Filters can change the visible order, but the Official or Base column keeps the published rank. Std. price is an estimate; hover over a value for its source and pricing details.

Columns

Choose the optional columns shown in this table.

Baseline filter leaderboard for DuelLab Benchmark
Rank Model Score Rank change Last tested Best Players
1Claude Opus 568.5Baseline2026-08-2191.014
2GPT-5.6 Sol55.8Baseline2026-08-2378.78
3Kimi K353.4Baseline2026-08-2398.98
4Claude Opus 4.549.0Baseline2026-08-0570.18
5Claude Opus 4.847.5Baseline2026-08-0573.87
6Grok 4.546.4Baseline2026-08-2385.58
7GPT-5.443.5Baseline2026-08-2275.28
8GLM-5.242.8Baseline2026-07-3159.616
9Qwen3.7 Max42.1Baseline2026-08-0568.97
10GPT-5.6 Terra41.9Baseline2026-08-0555.28
11GPT 4.141.5Baseline2026-08-0563.38
12Claude Sonnet 541.2Baseline2026-08-2357.18
13DeepSeek V4 Pro40.8Baseline2026-08-0566.67
14Gemini 3.5 Flash Lite40.6Baseline2026-08-2373.015
15Muse Spark 1.240.2Baseline2026-08-2359.48
16GPT-5.539.2Baseline2026-08-0554.07
17Laguna S 2.138.0Baseline2026-08-2257.912
18Gemini 3.6 Flash37.8Baseline2026-08-2355.016
19Minimax M337.6Baseline2026-08-0552.06
20Hy3 Preview37.2Baseline2026-08-0572.96
21GPT-5.4 Nano36.1Baseline2026-08-0556.112
22MiMo-V2.534.9Baseline2026-08-0550.817
23Nemotron 3 Ultra 550B A55B34.6Baseline2026-07-3159.77
24Inkling34.0Baseline2026-08-0569.012
25LongCat 2.033.9Baseline2026-08-0555.214
26GPT-5.4 Mini30.6Baseline2026-08-0543.88
27Ox Alpha30.6Baseline2026-08-2345.122
28DeepSeek V4 Flash29.2Baseline2026-08-0544.27
29Ling 3.0 Flash25.6Baseline2026-08-0540.610
30Qwen3.7 Plus25.2Baseline2026-08-0548.56
31GPT-5.6 Luna20.5Baseline2026-08-2337.18
32Gemini 3.1 Flash Lite16.3Baseline2026-07-3141.35
33GPT-OSS 120B14.5Baseline2026-08-2325.77
PGPT 4.1 Nano Provisional39.7Baseline2026-08-1439.71
PMiMo-V2.5-Pro Provisional39.1Baseline2026-08-0553.06
PQwen3.8 Max Provisional34.7Baseline2026-08-2365.38
PInkling Small Provisional31.9Baseline2026-08-2342.28
PNex N2 Pro Provisional31.4Baseline2026-08-0563.57
PGemma 4 31B Provisional31.1Baseline2026-08-0544.06
PStep 3.7 Flash Provisional31.0Baseline2026-08-2352.94
PMistral Medium 3.5 Provisional30.7Baseline2026-08-0540.94
PO3 Provisional22.8Baseline2026-08-2354.04
PNorth Mini Code Provisional20.4Baseline2026-08-0533.95

How this is scored

1. Generate players

Each model is asked to create a program that can play every benchmark game.

2. Play matches

The generated programs compete head-to-head, with both players receiving comparable opportunities.

3. Score each game

Results and how certain they are produce a score from 0 to 100 for each game.

4. Combine games

Game scores are combined into the main leaderboard score. A known failure to create a usable player contributes zero.

5. Keep settings clear

Every row uses the reasoning setting selected for this view.

6. Add context

Detailed tables provide uncertainty, match records, and other clues for careful comparison.