Benchmark menu

Current leaderboard score across public games (0–100)

Showing top 24 of 28 benchmarked models (updates when chart loads)

Scale: relative 0-100

  1. 79.0
    Grok 4.5
  2. 72.7
    Kimi K3
  3. 68.6
    GPT-5.6 Sol
  4. 67.2
    GPT-5.4
  5. 62.1
    GPT-5.6 Terra
  6. 59.3
    MiMo-V2.5-Pro
  7. 57.4
    O3
  8. 57.2
    Claude Opus 4.8
  9. 55.6
    GPT-5.5
  10. 55.4
    Claude Opus 5
  11. 55.1
    DeepSeek V4 Pro
  12. 53.4
    MiMo-V2.5
  13. 52.3
    Claude Opus 4.5
  14. 51.8
    DeepSeek V4 Flash
  15. 51.6
    Qwen3.7 Max
  16. 51.1
    Laguna S 2.1
  17. 49.6
    LongCat 2.0
  18. 49.1
    Gemma 4 31B
  19. 48.2
    GPT-5.4 Nano
  20. 46.6
    GPT-5.4 Mini
  21. 42.9
    Ling 3.0 Flash
  22. 42.3
    Qwen3.7 Plus
  23. 41.6
    GPT-5.6 Luna
  24. 39.5
    Hy3 Preview

Model leaderboard

One row per model; Best is the highest normalized score across that model's evaluated games in this reasoning variant view. Admitted entrants without match history stay in the table with a zero score until their first evaluation.

Reasoning level: Baseline Games: 8

Scroll sideways to see every column.

Baseline filter leaderboard for DuelLab Benchmark
Rank Model Score Rank movement Last measured Best Entries
1Grok 4.579.0Baseline2026-07-24100 × 28
2Kimi K372.7Baseline2026-07-241008
3GPT-5.6 Sol68.6Baseline2026-07-2698.78
4GPT-5.467.2Baseline2026-07-2692.18
5GPT-5.6 Terra62.1Baseline2026-07-2488.38
6MiMo-V2.5-Pro59.3Baseline2026-07-2677.36
7O357.4Baseline2026-07-2498.38
8Claude Opus 4.857.2Baseline2026-07-2668.87
9GPT-5.555.6Baseline2026-07-2677.07
10Claude Opus 555.4Baseline2026-07-2687.611
11DeepSeek V4 Pro55.1Baseline2026-07-261007
12MiMo-V2.553.4Baseline2026-07-2682.917
13Claude Opus 4.552.3Baseline2026-07-261008
14DeepSeek V4 Flash51.8Baseline2026-07-2670.518
15Qwen3.7 Max51.6Baseline2026-07-2688.87
16Laguna S 2.151.1Baseline2026-07-2660.912
17LongCat 2.049.6Baseline2026-07-2683.314
18Gemma 4 31B49.1Baseline2026-07-2468.76
19GPT-5.4 Nano48.2Baseline2026-07-2680.424
20GPT-5.4 Mini46.6Baseline2026-07-2484.724
21Ling 3.0 Flash42.9Baseline2026-07-2668.610
22Qwen3.7 Plus42.3Baseline2026-07-2496.16
23GPT-5.6 Luna41.6Baseline2026-07-2685.58
24Hy3 Preview39.5Baseline2026-07-2678.16
25Minimax M338.1Baseline2026-07-2668.66
26Inkling36.9Baseline2026-07-2679.112
27Step 3.7 Flash35.7Baseline2026-07-2381.97
28GPT-OSS 120B26.6Baseline2026-07-2478.18
PNex N2 Pro Provisional61.9Baseline2026-07-2488.37
PNorth Mini Code Provisional31.7Baseline2026-07-2660.45
PMistral Medium 3.5 Provisional26.4Baseline2026-07-2452.34

How this is scored

1. Generate players

Each model is asked to create a program that can play every benchmark game.

2. Play matches

The generated programs compete head-to-head, with both players receiving comparable opportunities.

3. Score each game

Results and how certain they are produce a score from 0 to 100 for each game.

4. Combine games

Game scores are combined into the main leaderboard score. A known failure to create a usable player contributes zero.

5. Keep settings clear

Every row uses the reasoning setting selected for this view.

6. Add context

Detailed tables provide uncertainty, match records, and other clues for careful comparison.