Benchmark menu

Every model with an official Overall rank is shown by default. Partial results keep their rank and show why their evidence is incomplete.

Charts

Charts

Use the Charts workspace to select multiple models, switch reasoning views, inspect per-game performance, and explore economics when that segment is available. Bookmark or share the current address to preserve the same model selection across visits.

All models as a static table

Compare two entries

Select two models to inspect their evidence together.

Applies to Scores, Reasoning variants, Head-to-head, Per game heatmap, Economics (when available), and Cycle detector.

Diagnostic ordering

Conservative score is the public leaderboard score after public-score adjustments; other buttons show diagnostic views when the selected models have that data.

Coming soon

Generation cost versus score will appear when the economics segment is available.

Overall min-max and mean; small dots are each game (averaged across reasoning variants where scores exist)
Reasoning variants
Combined view across tested provider settings.

Off-diagonal cells use win rate for the row model against the column opponent (W-L-D in the tooltip). Use Reasoning view to show one reasoning variant.

Strength heatmap (per-game scores averaged across reasoning variants)
Top selected models within each game
Show per-game comparison panels

Each panel ranks selected models within one game for the active reasoning view.

Non-transitive matchups

What it finds: Leaderboard averages can hide matchup loops. The detector searches the currently selected models and reasoning view for cases where A outperforms B, B outperforms C, and C outperforms A.

How to read it: In A → B → C → A, each arrow points from the model with the higher pair score to the model it outscored. Hover over a result for each edge’s win-loss-draw record and pair score.

Each arrow requires at least four matches and a pair score above 50%, with a draw counting as half a win. Results are ordered by the cycle’s weakest edge: weakest edge +x pp is the smallest of the three margins above 50%, while n is the total match count across all three edges. The detector shows up to ten cycles. These are descriptive matchup patterns, not tests of statistical significance. “No cycles” means none met these rules in the current selection and reasoning view; it does not prove that the matchups are transitive.

Static rankings — current page release

This table remains available without chart scripts. It describes this page's release; restoring a requested comparison and its filters requires JavaScript. Open benchmark rankings for the detailed tables.

Overall leaderboard scores

All officially ranked Overall models from the same data as the interactive charts. Partial rows keep their official rank and show that their evidence is incomplete.

Model Evidence status Rank Mean Min Max
GPT-6 AstraComplete evidence184.050.8100.0
Claude Fable 5.1Complete evidence280.733.6100.0
Claude Opus 5.5Complete evidence372.333.392.6
Claude Opus 5Complete evidence463.615.490.6
GLM 5.3Coverage incomplete: the official generation cohort is incomplete559.313.897.7
GPT-5.6 SolComplete evidence656.49.989.0
GPT-6 SolComplete evidence756.110.783.3
Grok 4.7Coverage incomplete: the official generation cohort is incomplete855.211.079.8
Claude Fable 5Complete evidence954.64.588.4
Grok 4.6Complete evidence1048.112.981.7
MiMo-V2.6-ProCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing; the official generation cohort is incomplete1147.44.577.5
GPT-5.4Complete evidence1246.52.1100.0
GLM 5.3 FlashComplete evidence1346.49.474.2
Gemini 3.8 FlashComplete evidence1446.08.683.6
Grok 4.5Complete evidence1545.40.079.3
Hy4 PreviewComplete evidence1645.310.581.4
Muse Spark 1.3Coverage incomplete: the official generation cohort is incomplete1745.312.877.7
Claude Opus 4.8Complete evidence1843.49.477.3
Qwen3.8 2.4t A95BComplete evidence1942.812.682.7
Gemini 3.6 FlashComplete evidence2042.46.584.8
Ox AlphaComplete evidence2142.113.366.9
GLM-5.3 FlashXComplete evidence2241.76.369.7
Gemini 3.5 FlashCoverage incomplete: fixed opponent-and-seat coverage is incomplete2340.53.283.1
MiMo-V2.6-FlashCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing2440.510.674.8
Qwen3.8 Max 0902Coverage incomplete: the official generation cohort is incomplete2540.50.478.0
GPT-5.5Complete evidence2640.310.683.9
Kimi K3Coverage incomplete: the official generation cohort is incomplete2740.110.788.2
GPT 4.1 NanoCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing; the official generation cohort is incomplete2840.140.140.1
Muse Spark 1.2Coverage incomplete: the official generation cohort is incomplete2939.710.374.0
DeepSeek V4.1 FlashComplete evidence3039.58.877.6
Qwen3.8 MaxComplete evidence3138.011.690.9
Claude Sonnet 5Complete evidence3237.613.364.3
Claude Opus 4.5Complete evidence3337.09.466.0
Qwen3.8 FlashCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing3436.78.373.0
GPT-5.6 TerraComplete evidence3535.38.375.5
Qwen3.8 27BCoverage incomplete: generation quality is below the publication threshold; fixed opponent-and-seat coverage is incomplete; the official generation cohort is incomplete3635.018.274.8
DeepSeek V4 ProComplete evidence3734.97.972.9
DeepSeek V4 Pro 0813Complete evidence3834.90.072.3
GLM-5.2Complete evidence3933.78.174.8
MiMo-V2.5-ProCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing4033.33.067.0
GPT 5Coverage incomplete: one or more documented settings are untested or ineligible; fixed opponent-and-seat coverage is incomplete4133.10.770.2
GPT-6 LunaCoverage incomplete: the official generation cohort is incomplete4232.411.278.8
Qwen3.7 MaxCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing4332.39.972.6
GPT 4.1Coverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing4430.76.456.5
GPT-5.6 LunaComplete evidence4530.54.278.9
O3Coverage incomplete: generation quality is below the publication threshold4630.47.453.0
Kimi K2.7 CodeCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing4730.210.152.1
GPT-5.4 MiniCoverage incomplete: one or more bracket scores are missing; one or more documented settings are untested or ineligible; generation-quality evidence is missing4829.30.059.8
Seed 2.0 CodeComplete evidence4929.37.874.7
GPT-5.4 NanoCoverage incomplete: one or more bracket scores are missing; one or more documented settings are untested or ineligible; generation-quality evidence is missing5027.99.857.8
DeepSeek V4 FlashComplete evidence5127.58.179.2
Laguna S 2.1Coverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing5227.36.948.8
Gemini 3.5 Flash LiteComplete evidence5326.89.354.8
Hy3 PreviewCoverage incomplete: one or more bracket scores are missing; one or more documented settings are untested or ineligible; generation-quality evidence is missing5426.53.961.5
Grok Build 0.1Coverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing5526.312.352.6
Qwen3.7 PlusCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing5625.99.071.6
LongCat 2.0Coverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing5725.17.242.2
GPT-OSS 120BComplete evidence5824.76.352.3
Step 3.7 FlashCoverage incomplete: generation quality is below the publication threshold5924.69.848.8
Muse Glimmer 30BComplete evidence6024.16.248.9
Gemma 4 31BCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing; the official generation cohort is incomplete6123.913.650.7
Nemotron 3 Ultra 550B A55BCoverage incomplete: one or more documented settings are untested or ineligible6223.34.954.9
Inkling SmallComplete evidence6323.33.657.1
Minimax M3Coverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing6423.10.041.7
InklingCoverage incomplete: generation quality is below the publication threshold6522.88.561.8
MiMo-V2.5Coverage incomplete: three distinct provider settings are not documented; one or more documented settings are untested or ineligible6622.36.949.2
Mistral Medium 3.5Coverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing6721.33.954.3
North Mini CodeCoverage incomplete: one or more bracket scores are missing; the provider capability record is unavailable; generation-quality evidence is missing; generation quality is below the publication threshold; fixed opponent-and-seat coverage is incomplete6820.613.749.7
Mercury 2.5 PreviewComplete evidence6919.79.839.9
Nex N2 ProCoverage incomplete: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing; generation quality is below the publication threshold7019.10.064.4
Ling 3.0 FlashCoverage incomplete: one or more bracket scores are missing; the provider capability record is unavailable; generation-quality evidence is missing; generation quality is below the publication threshold; fixed opponent-and-seat coverage is incomplete7116.64.038.8
Gemini 3.1 Flash LiteCoverage incomplete: one or more documented settings are untested or ineligible7216.40.831.7
Granite 4.2 8BCoverage incomplete: one or more bracket scores are missing; one or more documented settings are untested or ineligible; generation-quality evidence is missing; generation quality is below the publication threshold7310.51.122.5