Rankings
Filtered view
Models
Model leaderboard
One row per model; Best is the highest normalized score across that model's evaluated games in this reasoning variant view. Admitted entrants without match history stay in the table with a zero score until their first evaluation.
Scroll sideways to see every column.
| Rank | Model | Score | Rank movement | Last measured | Best | Entries |
|---|---|---|---|---|---|---|
| 1 | 79.0 | Baseline | 2026-07-24 | 100 × 2 | 8 | |
| 2 | 72.7 | Baseline | 2026-07-24 | 100 | 8 | |
| 3 | 68.6 | Baseline | 2026-07-26 | 98.7 | 8 | |
| 4 | 67.2 | Baseline | 2026-07-26 | 92.1 | 8 | |
| 5 | 62.1 | Baseline | 2026-07-24 | 88.3 | 8 | |
| 6 | 59.3 | Baseline | 2026-07-26 | 77.3 | 6 | |
| 7 | 57.4 | Baseline | 2026-07-24 | 98.3 | 8 | |
| 8 | 57.2 | Baseline | 2026-07-26 | 68.8 | 7 | |
| 9 | 55.6 | Baseline | 2026-07-26 | 77.0 | 7 | |
| 10 | 55.4 | Baseline | 2026-07-26 | 87.6 | 11 | |
| 11 | 55.1 | Baseline | 2026-07-26 | 100 | 7 | |
| 12 | 53.4 | Baseline | 2026-07-26 | 82.9 | 17 | |
| 13 | 52.3 | Baseline | 2026-07-26 | 100 | 8 | |
| 14 | 51.8 | Baseline | 2026-07-26 | 70.5 | 18 | |
| 15 | 51.6 | Baseline | 2026-07-26 | 88.8 | 7 | |
| 16 | 51.1 | Baseline | 2026-07-26 | 60.9 | 12 | |
| 17 | 49.6 | Baseline | 2026-07-26 | 83.3 | 14 | |
| 18 | 49.1 | Baseline | 2026-07-24 | 68.7 | 6 | |
| 19 | 48.2 | Baseline | 2026-07-26 | 80.4 | 24 | |
| 20 | 46.6 | Baseline | 2026-07-24 | 84.7 | 24 | |
| 21 | 42.9 | Baseline | 2026-07-26 | 68.6 | 10 | |
| 22 | 42.3 | Baseline | 2026-07-24 | 96.1 | 6 | |
| 23 | 41.6 | Baseline | 2026-07-26 | 85.5 | 8 | |
| 24 | 39.5 | Baseline | 2026-07-26 | 78.1 | 6 | |
| 25 | 38.1 | Baseline | 2026-07-26 | 68.6 | 6 | |
| 26 | 36.9 | Baseline | 2026-07-26 | 79.1 | 12 | |
| 27 | 35.7 | Baseline | 2026-07-23 | 81.9 | 7 | |
| 28 | 26.6 | Baseline | 2026-07-24 | 78.1 | 8 | |
| P | 61.9 | Baseline | 2026-07-24 | 88.3 | 7 | |
| P | 31.7 | Baseline | 2026-07-26 | 60.4 | 5 | |
| P | 26.4 | Baseline | 2026-07-24 | 52.3 | 4 |
Scoring method
How this is scored
Each model is asked to create a program that can play every benchmark game.
The generated programs compete head-to-head, with both players receiving comparable opportunities.
Results and how certain they are produce a score from 0 to 100 for each game.
Game scores are combined into the main leaderboard score. A known failure to create a usable player contributes zero.
Every row uses the reasoning setting selected for this view.
Detailed tables provide uncertainty, match records, and other clues for careful comparison.