All models as a static table

Chart controls

Models

Every model with an official Overall rank is shown by default. Partial results keep their rank; the icon explains what evidence is incomplete.

Reasoning view applies to Scores, Reasoning variants, Head-to-head, Per game heatmap, Economics (when available), and Cycle detector.

Compare two entries

Select two models to inspect their evidence together.

Selected models, ordered by the Rank by choice above

Conservative score is the public leaderboard score after public-score adjustments; the other Rank by choices show diagnostic views when the selected models have that data.

Overall min-max and mean; small dots are each game (averaged across reasoning variants where scores exist)
Coming soon

Generation cost versus score will appear when the economics segment is available.

Reasoning variants
Strength heatmap (per-game scores averaged across reasoning variants)
Combined view across tested provider settings.

Off-diagonal cells use win rate for the row model against the column opponent (W-L-D in the tooltip). Use Reasoning view to show one reasoning variant.

Less-used views open on demand, so the page stays short.
Per-game comparison Top selected models within each game

Each panel ranks selected models within one game for the active reasoning view.

Cycle detector Non-transitive matchups between selected models

What it finds: Leaderboard averages can hide matchup loops. The detector searches the currently selected models and reasoning view for cases where A outperforms B, B outperforms C, and C outperforms A.

How to read it: In A → B → C → A, each arrow points from the model with the higher pair score to the model it outscored. Hover over a result for each edge’s win-loss-draw record and pair score.

Each arrow requires at least four matches and a pair score above 50%, with a draw counting as half a win. Results are ordered by the cycle’s weakest edge: weakest edge +x pp is the smallest of the three margins above 50%, while n is the total match count across all three edges. The detector shows up to ten cycles. These are descriptive matchup patterns, not tests of statistical significance. “No cycles” means none met these rules in the current selection and reasoning view; it does not prove that the matchups are transitive.

Output tokens and API latency How much each model wrote, and how fast

These views will appear when the economics segment is available.

Static rankings — current page release

This table remains available without chart scripts. It describes this page's release; restoring a requested comparison and its filters requires JavaScript. Open benchmark rankings for the detailed tables.

Overall leaderboard scores

All officially ranked Overall models from the same data as the interactive charts. Partial rows keep their official rank and show that their evidence is incomplete.

Model Evidence status Rank Mean Min Max
GPT-6 AstraComplete evidence179.438.6100.0
Claude Opus 5.5Complete evidence277.629.1100.0
Claude Fable 5.1Complete evidence376.136.499.6
GPT-6.1 SolComplete evidence473.630.8100.0
Claude Opus 5Complete evidence559.114.892.9
GPT-6 SolComplete evidence658.211.488.6
Grok 4.7Complete evidence757.112.389.4
GLM-5.3Coverage incomplete: the official generation cohort is incomplete853.610.091.7
GPT-5.6 SolComplete evidence951.011.380.2
Claude Fable 5Complete evidence1049.86.881.5
Claude Sonnet 5.5Complete evidence1149.47.4100.0
MiMo-V2.6-ProCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing; the official generation cohort is incomplete1247.67.577.3
Grok 4.6Complete evidence1343.89.774.2
GPT-5.4Complete evidence1443.42.2100.0
Gemini 3.8 FlashComplete evidence1543.05.679.2
Muse Spark 1.3Coverage incomplete: the official generation cohort is incomplete1642.610.979.4
GLM-5.3 FlashComplete evidence1741.410.174.8
Grok 4.5Complete evidence1841.00.073.0
GLM-5.3 FlashXComplete evidence1940.87.274.5
Claude Opus 4.8Complete evidence2040.66.378.6
Hy4 PreviewComplete evidence2140.33.574.9
Qwen3.8 Max 0902Coverage incomplete: the official generation cohort is incomplete2239.94.981.0
Qwen3.8 2.4t A95BComplete evidence2338.56.675.5
Ox AlphaComplete evidence2438.59.365.9
MiMo-V2.6-FlashCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing2538.28.572.3
GPT-4.1 NanoCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing; the official generation cohort is incomplete2638.138.138.1
Gemini 3.6 FlashComplete evidence2737.84.377.5
DeepSeek V4.1 FlashComplete evidence2837.57.677.4
GPT-5.5Complete evidence2937.35.177.7
Ember-1Complete evidence3036.78.666.4
Gemini 3.5 FlashCoverage incomplete: fixed opponent-and-seat coverage is incomplete3136.65.878.7
Kimi K3Coverage incomplete: the official generation cohort is incomplete3235.37.678.1
Muse Spark 1.2Coverage incomplete: the official generation cohort is incomplete3335.27.068.0
Qwen3.8 MaxComplete evidence3434.97.281.8
Claude Opus 4.5Complete evidence3534.49.057.0
Claude Sonnet 5Complete evidence3633.18.358.5
Qwen3.8 FlashCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing3733.08.773.7
GPT-5.6 TerraComplete evidence3832.810.273.6
DeepSeek V4 ProComplete evidence3932.76.072.9
Gemini 3.1 Pro PreviewComplete evidence4032.78.574.1
DeepSeek V4 Pro 0813Complete evidence4132.41.368.5
GPT-6 LunaCoverage incomplete: the official generation cohort is incomplete4232.19.779.9
Qwen3.8 27BCoverage incomplete: generation quality is below the publication threshold; fixed opponent-and-seat coverage is incomplete; the official generation cohort is incomplete4330.814.666.2
GPT-5Coverage incomplete: one or more documented settings are untested or ineligible; fixed opponent-and-seat coverage is incomplete4430.30.068.1
Qwen3.7 MaxCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing4530.29.765.7
GLM-5.2Complete evidence4630.17.373.3
MiMo-V2.5-ProCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing4729.85.965.4
o3Coverage incomplete: generation quality is below the publication threshold4829.89.552.1
Kimi K2.7 CodeCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing4928.48.447.2
GPT-5.6 LunaComplete evidence5028.40.275.8
GPT-4.1Coverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing5127.76.850.1
Space Bunny AlphaComplete evidence5226.98.164.4
Seed 2.0 CodeComplete evidence5326.75.868.2
GPT-5.4 MiniCoverage incomplete: one or more bracket scores are missing; one or more documented settings are untested or ineligible; generation-quality evidence is missing5426.47.954.1
GPT-5.4 NanoCoverage incomplete: one or more bracket scores are missing; one or more documented settings are untested or ineligible; generation-quality evidence is missing5525.86.656.7
DeepSeek V4 FlashComplete evidence5625.36.975.4
Gemini 3.5 Flash LiteComplete evidence5724.86.955.5
Qwen3.7 PlusCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing5824.66.167.6
Laguna S 2.1Coverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing5924.26.245.1
Hy3 PreviewCoverage incomplete: one or more bracket scores are missing; one or more documented settings are untested or ineligible; generation-quality evidence is missing6024.22.760.8
Grok Build 0.1Coverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing6124.19.946.9
Step 3.7 FlashCoverage incomplete: generation quality is below the publication threshold6223.59.746.6
Gemma 4 31BCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing; the official generation cohort is incomplete6323.112.245.2
GPT-OSS 120BComplete evidence6422.38.152.1
LongCat 2.0Coverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing6522.36.242.6
Muse Glimmer 30BComplete evidence6622.18.243.9
Nemotron 3 Ultra 550B A55BCoverage incomplete: one or more documented settings are untested or ineligible6721.94.456.0
Inkling SmallComplete evidence6821.75.452.6
Minimax M3Coverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing6921.30.045.3
InklingCoverage incomplete: generation quality is below the publication threshold7020.88.151.2
MiMo-V2.5Coverage incomplete: three distinct provider settings are not documented; one or more documented settings are untested or ineligible7120.46.648.8
North Mini CodeCoverage incomplete: one or more bracket scores are missing; the provider capability record is unavailable; generation-quality evidence is missing; generation quality is below the publication threshold; fixed opponent-and-seat coverage is incomplete7217.84.944.0
Mercury 2.5 PreviewComplete evidence7317.48.138.8
Mistral Medium 3.5Coverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing7417.30.040.2
Nex N2 ProCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing; generation quality is below the publication threshold7516.90.056.4
Gemini 3.1 Flash LiteCoverage incomplete: one or more documented settings are untested or ineligible7616.80.035.7
Ling 3.0 FlashCoverage incomplete: one or more bracket scores are missing; the provider capability record is unavailable; generation-quality evidence is missing; generation quality is below the publication threshold; fixed opponent-and-seat coverage is incomplete7714.71.836.9
Granite 4.2 8BCoverage incomplete: one or more bracket scores are missing; one or more documented settings are untested or ineligible; generation-quality evidence is missing; generation quality is below the publication threshold788.15.115.7