Rankings
Filtered view
Models
Model leaderboard
One row per evaluated model; Best is the highest normalized score across that model's evaluated games in this reasoning variant view. Models appear after this reasoning variant has match history.
How to read
We ask each AI model to write a player program for every game. DuelLab runs those programs against each other. Higher scores mean stronger, better-supported performance against the current models.
Score is the normalized 0–100 result across public games. Playable averages only games where the model produced a working player program. Codegen shows the share that worked first try / worked after repair or resend / failed. Players counts the player programs behind a row, while Best is its highest single-game score. W / D / L shows rated wins, draws, and losses. The 90% band, Signal, and reasoning spread help show how much evidence supports the score and how much results vary.
Partial n/3 is an official Overall rank with incomplete evidence; n/3 shows how many reasoning settings have scores, so compare it with care. Provisional reasoning variants do not receive an official rank. Under-covered means evidence is sparse. Capped draws reached the shared move limit and still count as draws. Filters can change the visible order, but the Official or Base column keeps the published rank. Std. price is an estimate; hover over a value for its source and pricing details.
Columns
Choose the optional columns shown in this table.
Scoring method
How this is scored
Each model is asked to create a program that can play every benchmark game.
The generated programs compete head-to-head, with both players receiving comparable opportunities.
Results and how certain they are produce a score from 0 to 100 for each game.
Game scores are combined into the main leaderboard score. A known failure to create a usable player contributes zero.
Every row uses the reasoning setting selected for this view.
Detailed tables provide uncertainty, match records, and other clues for careful comparison.