Rankings
Filtered view
Models
Model leaderboard
One row per model; Best is the highest normalized score across that model's evaluated games in this reasoning variant view. Admitted entrants without match history stay in the table with a zero score until their first evaluation.
Scroll sideways to see every column.
| Rank | Model | Score | Rank movement | Last measured | Best | Entries |
|---|---|---|---|---|---|---|
| 1 | 70.5 | Baseline | 2026-07-26 | 99.3 | 14 | |
| 2 | 61.1 | Baseline | 2026-07-26 | 100 | 7 | |
| 3 | 59.4 | Baseline | 2026-07-26 | 81.3 | 23 | |
| 4 | 57.5 | Baseline | 2026-07-26 | 81.0 | 16 | |
| 5 | 52.1 | Baseline | 2026-07-26 | 77.2 | 11 | |
| 6 | 49.6 | Baseline | 2026-07-26 | 64.1 | 7 | |
| 7 | 45.6 | Baseline | 2026-07-26 | 66.0 | 8 | |
| 8 | 44.0 | Baseline | 2026-07-26 | 73.9 | 8 | |
| 9 | 41.3 | Baseline | 2026-07-26 | 90.0 | 8 | |
| 10 | 39.6 | Baseline | 2026-07-26 | 62.7 | 15 | |
| 11 | 39.5 | Baseline | 2026-07-24 | 58.6 | 7 | |
| 12 | 36.6 | Baseline | 2026-07-24 | 62.4 | 6 | |
| 13 | 34.7 | Baseline | 2026-07-24 | 49.2 | 2 | |
| 14 | 33.8 | Baseline | 2026-07-24 | 60.7 | 5 | |
| 15 | 33.3 | Baseline | 2026-07-24 | 47.4 | 2 | |
| 16 | 33.2 | Baseline | 2026-07-26 | 60.7 | 25 | |
| 17 | 32.5 | Baseline | 2026-07-24 | 75.0 | 8 | |
| 18 | 30.8 | Baseline | 2026-07-26 | 54.8 | 23 | |
| 19 | 30.5 | Baseline | 2026-07-24 | 52.5 | 5 | |
| 20 | 30.4 | Baseline | 2026-07-26 | 65.0 | 8 | |
| 21 | 29.9 | Baseline | 2026-07-26 | 48.3 | 8 | |
| 22 | 28.1 | Baseline | 2026-07-23 | 56.7 | 7 | |
| 23 | 25.2 | Baseline | 2026-07-26 | 34.6 | 7 | |
| 24 | 22.9 | Baseline | 2026-07-24 | 59.4 | 4 | |
| 25 | 22.7 | Baseline | 2026-07-24 | 55.5 | 8 | |
| 26 | 22.6 | Baseline | 2026-07-26 | 55.5 | 8 | |
| 27 | 22.5 | Baseline | 2026-07-24 | 30.8 | 9 | |
| 28 | 17.5 | Baseline | 2026-07-24 | 48.4 | 4 | |
| 29 | 13.5 | Baseline | 2026-07-24 | 36.4 | 4 | |
| 30 | 12.7 | Baseline | 2026-07-24 | 26.4 | 4 | |
| 31 | 9.9 | Baseline | 2026-07-24 | 20.0 | 2 | |
| 32 | 8.0 | Baseline | 2026-07-24 | 30.3 | 4 | |
| 33 | 0.0 | Baseline | 2026-07-24 | 0.0 | 2 | |
| P | 39.3 | Baseline | 2026-07-24 | 66.4 | 4 |
Scoring method
How this is scored
Each model is asked to create a program that can play every benchmark game.
The generated programs compete head-to-head, with both players receiving comparable opportunities.
Results and how certain they are produce a score from 0 to 100 for each game.
Game scores are combined into the main leaderboard score. A known failure to create a usable player contributes zero.
Every row uses the reasoning setting selected for this view.
Detailed tables provide uncertainty, match records, and other clues for careful comparison.