Benchmark menu

Normalized score across public games (0–100)

Showing top 24 of 48 benchmarked models (updates when chart loads)

Scale: relative 0-100

  1. 74.7
    Claude Fable 5
  2. 72.6
    Claude Opus 5
  3. 69.9
    GPT-5.6 Sol
  4. 57.0
    GPT-5.4
  5. 55.2
    Grok 4.5
  6. 53.9
    Claude Opus 4.8
  7. 53.3
    GPT-5.5
  8. 52.3
    Gemini 3.6 Flash
  9. 51.3
    Gemini 3.5 Flash
  10. 50.0
    Kimi K3
  11. 48.3
    Muse Spark 1.2
  12. 48.1
    Claude Opus 4.5
  13. 47.6
    GPT-5.6 Terra
  14. 46.4
    GLM-5.2
  15. 45.5
    Claude Sonnet 5
  16. 44.5
    GPT 5
  17. 43.5
    DeepSeek V4 Pro
  18. 42.4
    Kimi K2.7 Code
  19. 42.4
    Qwen3.7 Max
  20. 41.5
    GPT 4.1
  21. 41.1
    MiMo-V2.5-Pro
  22. 39.7
    GPT 4.1 Nano
  23. 39.5
    Grok Build 0.1
  24. 39.1
    Qwen3.8 Max
View

Overall leaderboard

How to read

We ask each AI model to write a player program for every game. DuelLab runs those programs against each other. Higher scores mean stronger, better-supported performance against the current models.

Score is the normalized 0–100 result across public games. Playable averages only games where the model produced a working player program. Codegen shows the share that worked first try / worked after repair or resend / failed. Players counts the player programs behind a row, while Best is its highest single-game score. W / D / L shows rated wins, draws, and losses. The 90% band, Signal, and reasoning spread help show how much evidence supports the score and how much results vary.

Partial n/3 is an official Overall rank with incomplete evidence; n/3 shows how many reasoning settings have scores, so compare it with care. Provisional reasoning variants do not receive an official rank. Under-covered means evidence is sparse. Capped draws reached the shared move limit and still count as draws. Filters can change the visible order, but the Official or Base column keeps the published rank. Std. price is an estimate; hover over a value for its source and pricing details.

Columns

Choose the optional columns shown in this table.

Overall Models leaderboard for DuelLab Benchmark
Rank Model Score W / D / L Levels Players Reasoning spread Signal
1Claude Fable 574.7791 / 129 / 208Partial 2/315XHigh 83.3 / Medium 66.1 / Low —
2Claude Opus 572.62180 / 473 / 542341XHigh 76.0 / Medium 73.5 / Thinking disabled 68.585.9u 14.1
3GPT-5.6 Sol69.91393 / 290 / 391324XHigh 87.3 / Medium 66.5 / None 55.890.1u 9.9
4GPT-5.457.01094 / 511 / 484323XHigh 72.7 / Medium 55.0 / None 43.591.9u 8.1
5Grok 4.555.21496 / 786 / 907324High 62.1 / Medium 57.2 / Low 46.497.2u 2.8
6Claude Opus 4.853.9839 / 484 / 473322XHigh 66.6 / Medium 47.4 / Thinking field omitted 47.588.8u 11.2
7GPT-5.553.3724 / 649 / 459322XHigh 67.5 / Medium 53.1 / None 39.289.3u 10.7
8Gemini 3.6 Flash52.32615 / 1589 / 1780342High 63.2 / Medium 55.8 / Minimal 37.899.4u 0.6
9Gemini 3.5 Flash51.3450 / 328 / 324Partial 2/314High 47.3 / Medium 55.4 / Minimal —
10Kimi K350.01319 / 991 / 919Partial 3/321Max 48.9 / High 47.6 / Low 53.499.3u 0.7
11Muse Spark 1.248.3621 / 317 / 525Partial 3/317XHigh 52.0 / Medium 52.8 / Minimal 40.291.1u 8.9
12Claude Opus 4.548.11274 / 912 / 1023Partial 3/324High effort + 32,000 thinking tokens 43.0 / Medium effort + 16,000 thinking tokens 52.3 / Thinking field omitted 49.096.1u 3.9
13GPT-5.6 Terra47.6698 / 661 / 603324XHigh 57.5 / Medium 43.4 / None 41.988.2u 11.8
14GLM-5.246.4581 / 601 / 548Partial 3/326XHigh 53.3 / Medium 43.2 / Reasoning disabled 42.878.2u 21.8
15Claude Sonnet 545.5748 / 550 / 736324XHigh 54.3 / Medium 40.9 / Thinking disabled 41.289.5u 10.5
16GPT 544.5417 / 349 / 434Partial 2/315High 46.2 / Medium 42.8 / Minimal —
17DeepSeek V4 Pro43.5388 / 332 / 388Partial 2/313XHigh 46.3 / High — / Reasoning disabled 40.8
18Kimi K2.7 Code42.4216 / 163 / 193Partial 1/36Intensive track (legacy evidence incomplete) — / Reasoning enabled 42.4 / Baseline track (legacy evidence incomplete) —
19Qwen3.7 Max42.4393 / 430 / 397Partial 2/314Intensive track (legacy evidence incomplete) — / Reasoning enabled 42.7 / Reasoning disabled 42.1
20GPT 4.141.5199 / 228 / 224Partial 1/38Intensive track (legacy evidence incomplete) — / Balanced track (legacy evidence incomplete) — / None 41.5
21MiMo-V2.5-Pro41.1367 / 232 / 369Partial 2/311Intensive track (legacy evidence incomplete) — / Reasoning enabled 43.0 / Reasoning disabled 39.1
22GPT 4.1 Nano39.760 / 35 / 67Partial 1/31Intensive track (legacy evidence incomplete) — / Balanced track (legacy evidence incomplete) — / Reasoning field omitted 39.7
23Grok Build 0.139.5209 / 276 / 241Partial 1/38Intensive track (legacy evidence incomplete) — / Reasoning enabled 39.5 / Baseline track (legacy evidence incomplete) —
24Qwen3.8 Max39.1162 / 132 / 115Partial 3/323XHigh 42.4 / Medium 40.1 / Minimal 34.728.3u 71.7
25GPT-5.4 Mini38.7407 / 517 / 457Partial 2/316XHigh — / Medium 46.8 / None 30.6
26GPT-5.4 Nano37.9731 / 507 / 802Partial 2/324XHigh — / Medium 39.7 / None 36.1
27GPT-5.6 Luna37.8597 / 604 / 857324XHigh 52.8 / Medium 40.2 / None 20.589.8u 10.2
28O337.51649 / 597 / 1760Partial 3/329High 42.8 / Medium 46.7 / Low 22.897.5u 2.5
29Ox Alpha36.3369 / 224 / 369366Max 44.3 / High 33.9 / Low 30.616.2u 83.8
30Hy3 Preview36.1309 / 346 / 369Partial 2/313High 35.0 / Low — / None 37.2
31Gemma 4 31B34.6310 / 234 / 406Partial 2/310Intensive track (legacy evidence incomplete) — / Reasoning enabled 38.0 / Reasoning disabled 31.1
32Qwen3.7 Plus34.4312 / 336 / 423Partial 2/312Intensive track (legacy evidence incomplete) — / Reasoning enabled 43.5 / Reasoning disabled 25.2
33Gemini 3.5 Flash Lite34.41581 / 1845 / 2652342High 35.4 / Medium 27.0 / Minimal 40.699.3u 0.7
34DeepSeek V4 Flash33.3494 / 689 / 791Partial 3/327XHigh 38.7 / Medium 32.1 / Reasoning disabled 29.286.9u 13.1
35Laguna S 2.133.2687 / 1114 / 1117Partial 2/323Intensive track (legacy evidence incomplete) — / Enabled 28.3 / None 38.0
36Inkling Small33.0100 / 148 / 156Partial 3/323Max 35.2 / Medium 31.9 / None 31.926.7u 73.3
37Minimax M332.8347 / 330 / 505Partial 2/313Intensive track (legacy evidence incomplete) — / Reasoning enabled 28.0 / Reasoning disabled 37.6
38MiMo-V2.532.7832 / 1059 / 1095Partial 3/337High 31.0 / Reasoning enabled 32.3 / Reasoning disabled 34.986.2u 13.8
39LongCat 2.032.5836 / 896 / 1363Partial 2/326Intensive track (legacy evidence incomplete) — / Enabled 31.1 / None 33.9
40Nemotron 3 Ultra 550B A55B31.9383 / 614 / 592Partial 3/321High 31.3 / Medium 29.8 / Reasoning disabled 34.688.4u 11.6
41Step 3.7 Flash30.8425 / 318 / 682Partial 3/313High 29.1 / Medium 32.3 / Low 31.095.6u 4.4
42Nex N2 Pro29.6257 / 312 / 431Partial 2/312Intensive track (legacy evidence incomplete) — / Reasoning enabled 27.8 / Reasoning disabled 31.4
43Inkling28.7746 / 971 / 1496Partial 3/327Max 27.4 / Medium 24.7 / None 34.096.8u 3.2
44GPT-OSS 120B28.3378 / 665 / 1040319High 34.8 / Medium 35.6 / Low 14.593.2u 6.8
45North Mini Code27.0246 / 352 / 460Partial 2/317Intensive track (legacy evidence incomplete) — / Enabled 33.6 / None 20.4
46Gemini 3.1 Flash Lite25.5302 / 341 / 704Partial 3/316High 30.6 / Medium 29.5 / Reasoning disabled 16.389.0u 11.0
47Mistral Medium 3.524.2192 / 121 / 381Partial 2/38High 17.8 / Balanced track (legacy evidence incomplete) — / None 30.7
48Ling 3.0 Flash23.7418 / 861 / 901Partial 2/322Intensive track (legacy evidence incomplete) — / Enabled 21.8 / None 25.6

How this is scored

1. Generate players

Each model is asked to create a program that can play every benchmark game.

2. Play matches

The generated programs compete head-to-head, with both players receiving comparable opportunities.

3. Score each game

Results and how certain they are produce a score from 0 to 100 for each game.

4. Combine games

Game scores are combined into the main leaderboard score. A known failure to create a usable player contributes zero.

5. Keep settings clear

The reasoning settings are combined into one row for each model.

6. Add context

Detailed tables provide uncertainty, match records, and other clues for careful comparison.