Benchmark menu

Current leaderboard score across public games (0–100)

Showing top 24 of 24 benchmarked models (updates when chart loads)

Scale: relative 0-100

  1. 66.9
    Claude Fable 5
  2. 64.9
    Claude Opus 5
  3. 58.6
    GPT-5.4
  4. 55.2
    GPT-5.6 Sol
  5. 41.9
    GPT-5.5
  6. 41.4
    GPT-5.6 Terra
  7. 37.9
    Claude Opus 4.5
  8. 37.2
    Claude Sonnet 5
  9. 34.3
    Kimi K3
  10. 34.0
    DeepSeek V4 Pro
  11. 32.4
    GPT-5.6 Luna
  12. 30.1
    GPT-5.4 Mini
  13. 28.2
    GPT-5.4 Nano
  14. 25.8
    GLM-5.2
  15. 20.3
    Gemini 3.5 Flash
  16. 19.5
    GPT 5
  17. 19.4
    Claude Opus 4.8
  18. 19.3
    DeepSeek V4 Flash
  19. 15.6
    Grok 4.5
  20. 14.1
    Hy3 Preview
  21. 7.0
    Nemotron 3 Ultra 550B A55B
  22. 1.8
    Gemini 3.1 Flash Lite
  23. 0.9
    GPT-OSS 120B
  24. 0.6
    Step 3.7 Flash

Model leaderboard

One row per model; Best is the highest normalized score across that model's evaluated games in this reasoning variant view. Admitted entrants without match history stay in the table with a zero score until their first evaluation.

Reasoning level: Intensive Games: 8

Scroll sideways to see every column.

Intensive filter leaderboard for DuelLab Benchmark
Rank Model Score Rank movement Last measured Best Entries
1Claude Fable 566.9Baseline2026-07-2691.38
2Claude Opus 564.9Baseline2026-07-2695.213
3GPT-5.458.6Baseline2026-07-261007
4GPT-5.6 Sol55.2Baseline2026-07-2699.48
5GPT-5.541.9Baseline2026-07-261007
6GPT-5.6 Terra41.4Baseline2026-07-2662.38
7Claude Opus 4.537.9Baseline2026-07-2682.46
8Claude Sonnet 537.2Baseline2026-07-2664.210
9Kimi K334.3Baseline2026-07-2442.18
10DeepSeek V4 Pro34.0Baseline2026-07-2654.56
11GPT-5.6 Luna32.4Baseline2026-07-2448.28
12GPT-5.4 Mini30.1Baseline2026-07-2641.912
13GPT-5.4 Nano28.2Baseline2026-07-2629.114
14GLM-5.225.8Baseline2026-07-2641.218
15Gemini 3.5 Flash20.3Baseline2026-07-2424.72
16GPT 519.5Baseline2026-07-2442.74
17Claude Opus 4.819.4Baseline2026-07-2620.17
18DeepSeek V4 Flash19.3Baseline2026-07-2436.111
19Grok 4.515.6Baseline2026-07-2643.112
20Hy3 Preview14.1Baseline2026-07-2422.72
21Nemotron 3 Ultra 550B A55B7.0Baseline2026-07-248.24
22Gemini 3.1 Flash Lite1.8Baseline2026-07-242.01
23GPT-OSS 120B0.9Baseline2026-07-242.02
24Step 3.7 Flash0.6Baseline2026-07-241.42
PO3 Provisional10.0Baseline2026-07-2410.01
PMistral Medium 3.5 Provisional3.6Baseline2026-07-247.82

How this is scored

1. Generate players

Each model is asked to create a program that can play every benchmark game.

2. Play matches

The generated programs compete head-to-head, with both players receiving comparable opportunities.

3. Score each game

Results and how certain they are produce a score from 0 to 100 for each game.

4. Combine games

Game scores are combined into the main leaderboard score. A known failure to create a usable player contributes zero.

5. Keep settings clear

Every row uses the reasoning setting selected for this view.

6. Add context

Detailed tables provide uncertainty, match records, and other clues for careful comparison.