Rankings
Filtered view
Models
Model leaderboard
One row per model; Best is the highest normalized score across that model's evaluated games in this reasoning variant view. Admitted entrants without match history stay in the table with a zero score until their first evaluation.
Scroll sideways to see every column.
| Rank | Model | Score | Rank movement | Last measured | Best | Entries |
|---|---|---|---|---|---|---|
| 1 | 66.9 | Baseline | 2026-07-26 | 91.3 | 8 | |
| 2 | 64.9 | Baseline | 2026-07-26 | 95.2 | 13 | |
| 3 | 58.6 | Baseline | 2026-07-26 | 100 | 7 | |
| 4 | 55.2 | Baseline | 2026-07-26 | 99.4 | 8 | |
| 5 | 41.9 | Baseline | 2026-07-26 | 100 | 7 | |
| 6 | 41.4 | Baseline | 2026-07-26 | 62.3 | 8 | |
| 7 | 37.9 | Baseline | 2026-07-26 | 82.4 | 6 | |
| 8 | 37.2 | Baseline | 2026-07-26 | 64.2 | 10 | |
| 9 | 34.3 | Baseline | 2026-07-24 | 42.1 | 8 | |
| 10 | 34.0 | Baseline | 2026-07-26 | 54.5 | 6 | |
| 11 | 32.4 | Baseline | 2026-07-24 | 48.2 | 8 | |
| 12 | 30.1 | Baseline | 2026-07-26 | 41.9 | 12 | |
| 13 | 28.2 | Baseline | 2026-07-26 | 29.1 | 14 | |
| 14 | 25.8 | Baseline | 2026-07-26 | 41.2 | 18 | |
| 15 | 20.3 | Baseline | 2026-07-24 | 24.7 | 2 | |
| 16 | 19.5 | Baseline | 2026-07-24 | 42.7 | 4 | |
| 17 | 19.4 | Baseline | 2026-07-26 | 20.1 | 7 | |
| 18 | 19.3 | Baseline | 2026-07-24 | 36.1 | 11 | |
| 19 | 15.6 | Baseline | 2026-07-26 | 43.1 | 12 | |
| 20 | 14.1 | Baseline | 2026-07-24 | 22.7 | 2 | |
| 21 | 7.0 | Baseline | 2026-07-24 | 8.2 | 4 | |
| 22 | 1.8 | Baseline | 2026-07-24 | 2.0 | 1 | |
| 23 | 0.9 | Baseline | 2026-07-24 | 2.0 | 2 | |
| 24 | 0.6 | Baseline | 2026-07-24 | 1.4 | 2 | |
| P | 10.0 | Baseline | 2026-07-24 | 10.0 | 1 | |
| P | 3.6 | Baseline | 2026-07-24 | 7.8 | 2 |
Scoring method
How this is scored
Each model is asked to create a program that can play every benchmark game.
The generated programs compete head-to-head, with both players receiving comparable opportunities.
Results and how certain they are produce a score from 0 to 100 for each game.
Game scores are combined into the main leaderboard score. A known failure to create a usable player contributes zero.
Every row uses the reasoning setting selected for this view.
Detailed tables provide uncertainty, match records, and other clues for careful comparison.