Benchmark menu

Normalized score across public games (0–100)

Showing top 24 of 98 benchmarked reasoning variants (updates when chart loads)

Scale: relative 0-100

  1. 84.6
    Claude Fable 5
  2. 76.7
    Claude Opus 5
  3. 72.4
    Claude Opus 5
  4. 72.0
    GPT-5.4
  5. 67.3
    Claude Opus 5
  6. 66.4
    Claude Fable 5
  7. 66.3
    GPT-5.6 Sol
  8. 64.9
    GPT-5.5
  9. 64.9
    Claude Opus 4.8
  10. 57.0
    GPT-5.6 Sol
  11. 55.5
    GPT-5.6 Terra
  12. 53.5
    Gemini 3.5 Flash
  13. 53.5
    Grok 4.5
  14. 51.3
    Gemini 3.6 Flash
  15. 50.7
    GPT-5.4
  16. 50.6
    GPT-5.5
  17. 50.3
    GPT-5.6 Luna
  18. 48.0
    GPT-5.4
  19. 47.1
    Claude Opus 4.5
  20. 47.1
    Kimi K3
  21. 47.0
    Grok 4.5
  22. 46.3
    Gemini 3.6 Flash
  23. 45.7
    Claude Opus 4.5
  24. 45.0
    Claude Opus 4.8
View
Reasoning

Detailed leaderboard

How to read

We ask each AI model to write a player program for every game. DuelLab runs those programs against each other. Higher scores mean stronger, better-supported performance against the current models.

Score is the normalized 0–100 result across public games. Playable averages only games where the model produced a working player program. Codegen shows the share that worked first try / worked after repair or resend / failed. Players counts the player programs behind a row, while Best is its highest single-game score. W / D / L shows rated wins, draws, and losses. The 90% band, Signal, and reasoning spread help show how much evidence supports the score and how much results vary.

Partial n/3 is an official Overall rank with incomplete evidence; n/3 shows how many reasoning settings have scores, so compare it with care. Provisional reasoning variants do not receive an official rank. Under-covered means evidence is sparse. Capped draws reached the shared move limit and still count as draws. Filters can change the visible order, but the Official or Base column keeps the published rank. Std. price is an estimate; hover over a value for its source and pricing details.

Columns

Choose the optional columns shown in this table.

Reasoning Variants leaderboard for DuelLab Benchmark
Rank Model Reasoning Score Playable W / D / L Best Std. price Codegen By game
1Claude Fable 5XHigh84.684.6502 / 94 / 116100 × 2$1.58List-price estimate63%37%0%n 8
2Claude Opus 5XHigh76.780.8674 / 63 / 157100$1.38List-price estimate81%0%19%n 16
3Claude Opus 5Medium72.474.9905 / 176 / 23395.7$0.45List-price estimate88%0%12%n 16
4GPT-5.4XHigh72.072.0441 / 136 / 119100 × 2$2.64List-price estimate88%12%0%n 8
5Claude Opus 5None67.369.6841 / 302 / 24699.2$0.29List-price estimate88%0%12%n 16
6Claude Fable 5Medium66.468.7405 / 68 / 14784.0$0.50List-price estimate88%0%12%n 8
7GPT-5.6 SolXHigh66.366.3420 / 79 / 201100$10.06List-price estimate100%0%0%n 8
8GPT-5.5XHigh64.967.1368 / 124 / 14494.1$2.74List-price estimate88%0%12%n 8
9Claude Opus 4.8XHigh64.967.1394 / 83 / 18185.7$0.93List-price estimate75%13%12%n 8
10GPT-5.6 SolMedium57.058.7380 / 154 / 18085.1$0.47List-price estimate78%11%11%n 9
11GPT-5.6 TerraXHigh55.555.5325 / 152 / 23179.0$8.17List-price estimate63%37%0%n 8
12Gemini 3.5 FlashMedium53.553.5311 / 234 / 18793.2$0.35Provider-reported0%100%0%n 8
13Grok 4.5Medium53.553.5395 / 192 / 24780.5$0.06Provider-reported88%12%0%n 8
14Gemini 3.6 FlashMedium51.353.1516 / 306 / 23291.7$0.26Provider-reported88%0%12%n 16
15GPT-5.4None50.750.7292 / 262 / 154100.0$0.09List-price estimate75%25%0%n 8
16GPT-5.5Medium50.650.6299 / 253 / 22268.9$0.44List-price estimate88%12%0%n 8
17GPT-5.6 LunaXHigh50.351.8308 / 156 / 23481.1$3.16List-price estimate78%11%11%n 9
18GPT-5.4Medium48.048.0305 / 235 / 20095.4$0.50List-price estimate88%12%0%n 8
19Claude Opus 4.5None47.147.1282 / 267 / 21368.2$0.12List-price estimate100%0%0%n 8
20Kimi K3Low47.147.1308 / 221 / 17168.6$0.08Provider-reported88%12%0%n 8
21Grok 4.5High47.047.0296 / 123 / 18366.9$0.07Provider-reported75%25%0%n 8
22Gemini 3.6 FlashHigh46.351.3293 / 95 / 14574.5$0.46Provider-reported34%33%33%n 12
23Claude Opus 4.5Medium45.745.7295 / 195 / 22369.0$0.18List-price estimate100%0%0%n 8
24Claude Opus 4.8None45.046.5264 / 157 / 18766.5$0.16List-price estimate88%0%12%n 8
25GPT-5.4 MiniMedium44.644.6308 / 250 / 24762.1$0.25List-price estimate88%12%0%n 8
26Claude Opus 4.5Historical43.544.2350 / 388 / 35660.0$0.26List-price estimate88%6%6%n 16
27Claude Opus 4.8Medium43.443.4258 / 297 / 20383.1$0.22List-price estimate100%0%0%n 8
28GPT-5.6 LunaMedium43.044.3246 / 295 / 22970.7$0.04List-price estimate89%0%11%n 9
29GPT-5.6 SolNone42.844.1272 / 189 / 23562.4$0.12List-price estimate89%0%11%n 9
30O3Medium42.645.8247 / 157 / 22265.5$0.05List-price estimate75%0%25%n 8
31DeepSeek V4 ProXHigh41.244.3202 / 132 / 24271.1$0.08Provider-reported13%62%25%n 8
32GPT-5.6 TerraMedium40.640.6247 / 287 / 25054.9$0.17List-price estimate63%37%0%n 8
33GLM-5.2None40.640.6365 / 376 / 39964.3$0.02Provider-reported56%44%0%n 16
34Qwen3.7 PlusEnabled40.543.5222 / 198 / 21674.2$0.05Provider-reported13%62%25%n 8
35GLM-5.2Medium40.440.4177 / 225 / 16069.7$0.07Provider-reported0%100%0%n 8
36Claude Sonnet 5Medium40.241.5275 / 237 / 21252.8$0.08List-price estimate75%13%12%n 8
37GPT-5.6 TerraNone39.640.8212 / 298 / 24052.2$0.07List-price estimate78%11%11%n 9
38Grok 4.5Low39.439.4220 / 182 / 16462.9$0.07Provider-reported88%12%0%n 8
39DeepSeek V4 ProNone39.240.6216 / 232 / 21268.9$0.02Provider-reported50%38%12%n 8
40GPT 5Medium38.639.9217 / 173 / 24863.5$0.18List-price estimate75%13%12%n 8
41Qwen3.7 MaxNone38.639.9204 / 211 / 22366.1$0.02Provider-reported88%0%12%n 8
42MiMo-V2.5-ProNone38.139.6213 / 136 / 19951.2$0.03Provider-reported57%29%14%n 7
43Kimi K3Max38.041.3145 / 101 / 10252.8$0.53Provider-reported14%57%29%n 7
44GPT-5.4 NanoMedium37.940.7445 / 221 / 49265.7$0.02List-price estimate31%44%25%n 16
45GPT 5High37.837.8241 / 225 / 26868.2$0.28List-price estimate75%25%0%n 8
46Gemini 3.5 FlashHigh37.740.5179 / 148 / 19377.6$0.49Provider-reported0%75%25%n 8
47MiMo-V2.5-ProEnabled37.540.3184 / 125 / 22169.2$0.13Provider-reported0%75%25%n 8
48GPT 4.1None37.437.4210 / 257 / 26260.7$0.06List-price estimate63%37%0%n 8
49Qwen3.7 MaxEnabled37.338.6223 / 263 / 24463.1$0.09Provider-reported13%75%12%n 8
50Kimi K2.7 CodeEnabled37.340.1240 / 183 / 22153.7$0.16Provider-reported0%75%25%n 8
51Kimi K3High37.137.1188 / 172 / 18258.5$0.29Provider-reported88%12%0%n 8
52Laguna S 2.1None36.838.2382 / 415 / 41552.6$0.0030Provider-reported50%36%14%n 14
53DeepSeek V4 FlashXHigh36.739.4170 / 149 / 21769.7$0.0100Provider-reported38%37%25%n 8
54Gemini 3.5 Flash LiteMinimal36.537.1299 / 390 / 33060.7$0.02Provider-reported56%38%6%n 16
55GPT-5.5None36.137.4155 / 344 / 17948.4$0.14List-price estimate88%0%12%n 8
56Gemini 3.6 FlashMinimal36.136.1336 / 473 / 34449.2$0.06Provider-reported75%25%0%n 16
57GPT-5.4 NanoNone35.738.3353 / 331 / 43057.2$0.0092List-price estimate31%44%25%n 16
58Minimax M3None35.337.9223 / 133 / 22451.4$0.03Provider-reported50%25%25%n 8
59O3High35.240.1546 / 122 / 51551.0$0.09Combined: Recorded attempt + Standard list-price estimate59%0%41%n 32
60MiMo-V2.5Reasoning34.337.7299 / 281 / 32255.2$0.0024Provider-reported44%25%31%n 16
61Grok Build 0.1Enabled34.034.0221 / 315 / 29051.8$0.06Provider-reported0%100%0%n 8
62Hy3 PreviewNone33.536.0146 / 238 / 15071.4$0.0013Provider-reported50%25%25%n 8
63MiMo-V2.5Reasoning33.335.7183 / 129 / 25855.7$0.0038Provider-reported63%12%25%n 8
64Claude Opus 4.5High33.233.2198 / 153 / 24955.4$0.59List-price estimate100%0%0%n 8
65LongCat 2.0None33.034.1371 / 512 / 52151.8$0.01Provider-reported50%38%12%n 16
66InklingNone33.035.4335 / 380 / 49163.3$0.04Provider-reported44%31%25%n 16
67Gemma 4 31BEnabled32.437.2168 / 74 / 19445.9$0.0066Provider-reported0%57%43%n 7
68GPT-OSS 120BMedium32.233.3139 / 312 / 23842.1$0.0025Provider-reported13%75%12%n 8
69Nemotron 3 Ultra 550B A55BNone31.833.8135 / 172 / 20553.8$0.03Provider-reported56%22%22%n 9
70LongCat 2.0Enabled31.134.9160 / 112 / 19852.2$0.10Provider-reported64%0%36%n 11
71Nemotron 3 Ultra 550B A55BMedium30.735.5119 / 182 / 17766.1$0.04Provider-reported0%56%44%n 9
72DeepSeek V4 FlashMedium30.531.6241 / 360 / 42143.0$0.0032Provider-reported63%25%12%n 16
73Gemini 3.6 FlashHigh30.530.592 / 51 / 7951.6$0.55Provider-reported100%0%0%n 4
74MiMo-V2.5High30.434.247 / 224 / 8162.6$0.02Provider-reported0%63%37%n 8
75Gemini 3.5 Flash LiteHigh30.031.0227 / 180 / 33147.7$0.06Provider-reported81%6%13%n 16
76MiMo-V2.5Reasoning29.731.6225 / 220 / 31537.4$0.02Provider-reported0%78%22%n 9
77Gemma 4 31BNone29.430.6163 / 186 / 25344.5$0.0024Provider-reported72%14%14%n 7
78Hy3 PreviewHigh28.830.6194 / 153 / 28554.1$0.0032Provider-reported56%22%22%n 9
79GPT-5.4 MiniNone28.628.6127 / 336 / 26139.7$0.02List-price estimate75%25%0%n 8
80North Mini CodeEnabled28.530.6185 / 186 / 26143.6$0.00Provider-reported38%37%25%n 8
81GPT-OSS 120BHigh27.831.398 / 166 / 15066.6$0.0088Provider-reported0%63%37%n 8
82Laguna S 2.1Enabled27.128.8174 / 366 / 29841.3$0.0034Provider-reported50%29%21%n 14
83Gemini 3.1 Flash LiteMedium26.929.0159 / 160 / 32349.4$0.01Provider-reported25%50%25%n 8
84Gemini 3.5 Flash LiteMedium26.727.6198 / 378 / 41145.5$0.04Provider-reported88%0%12%n 16
85Step 3.7 FlashHigh26.029.3144 / 74 / 25542.1$0.06Provider-reported0%63%37%n 8
86MiMo-V2.5Medium25.828.5142 / 335 / 27148.6$0.02Provider-reported9%58%33%n 12
87DeepSeek V4 FlashNone25.226.1134 / 248 / 25636.2$0.0023Provider-reported63%25%12%n 8
88Minimax M3Enabled24.925.8147 / 240 / 34546.3$0.07Provider-reported0%88%12%n 8
89Qwen3.7 PlusNone24.727.3116 / 175 / 25242.8$0.0087Provider-reported45%22%33%n 9
90Ling 3.0 FlashNone24.627.6203 / 662 / 43340.2$0.00Provider-reported38%25%37%n 16
91GPT-OSS 120BLow24.225.1110 / 127 / 29444.4$0.0016Provider-reported63%25%12%n 8
92LongCat 2.0Enabled23.923.983 / 71 / 17644.4$0.07Provider-reported40%60%0%n 5
93Nemotron 3 Ultra 550B A55BHigh23.827.4158 / 323 / 30239.9$0.03Provider-reported12%44%44%n 16
94InklingMax23.627.2123 / 167 / 27654.7$0.26Provider-reported25%31%44%n 16
95Gemini 3.1 Flash LiteHigh22.825.793 / 128 / 24837.3$0.03Provider-reported0%63%37%n 8
96Nex N2 ProEnabled21.724.4126 / 56 / 30239.5$0.10Provider-reported0%63%37%n 8
97GPT-5.6 LunaNone19.919.976 / 302 / 33537.0$0.02List-price estimate88%12%0%n 8
98Gemini 3.1 Flash LiteNone17.019.273 / 85 / 20846.3$0.0059Provider-reported38%25%37%n 8
PClaude Sonnet 5 ProvisionalXHigh52.252.256 / 70 / 4852.2$0.47List-price estimate100%0%0%n 2
PGPT 4.1 Nano ProvisionalReasoning47.847.834 / 26 / 2247.8$0.0023Combined: Recorded attempt + Standard list-price estimate100%0%0%n 1
PClaude Sonnet 5 ProvisionalNone44.844.88 / 66 / 1044.8$0.09Combined: Provider-reported + Recorded attempt100%0%0%n 1
PGLM-5.2 ProvisionalXHigh42.242.227 / 21 / 1842.2$0.09Provider-reported100%0%0%n 2
PStep 3.7 Flash ProvisionalLow31.737.7109 / 55 / 12461.2$0.02Provider-reported38%12%50%n 8
PMistral Medium 3.5 ProvisionalNone29.735.4124 / 107 / 17044.6$0.04Provider-reported25%25%50%n 8
PStep 3.7 Flash ProvisionalMedium29.234.797 / 149 / 16041.8$0.04Provider-reported0%50%50%n 8
PNex N2 Pro ProvisionalNone28.134.5154 / 291 / 19359.0$0.16Provider-reported0%44%56%n 16
PO3 ProvisionalLow27.733.062 / 104 / 13444.1$0.11Combined: Recorded attempt + Standard list-price estimate50%0%50%n 8
PInkling ProvisionalMedium24.431.299 / 111 / 19354.9$0.08Provider-reported31%6%63%n 16
PLing 3.0 Flash ProvisionalEnabled21.226.7123 / 132 / 23740.7$0.00Provider-reported7%33%60%n 15
PNorth Mini Code ProvisionalNone19.325.887 / 192 / 25126.3$0.00Provider-reported19%12%69%n 16
PMistral Medium 3.5 ProvisionalHigh17.621.580 / 19 / 25732.0$0.19Provider-reported22%22%56%n 9

How this is scored

1. Generate players

Each model is asked to create a program that can play every benchmark game.

2. Play matches

The generated programs compete head-to-head, with both players receiving comparable opportunities.

3. Score each game

Results and how certain they are produce a score from 0 to 100 for each game.

4. Combine games

Game scores are combined into the main leaderboard score. A known failure to create a usable player contributes zero.

5. Keep settings clear

Each model and reasoning setting keeps its own row in the standings.

6. Add context

Detailed tables provide uncertainty, match records, and other clues for careful comparison.