Reasoning view
Applies to Scores, Reasoning variants, Head-to-head, Per game heatmap, Economics (when available), and Cycle detector.
Rank by
Diagnostic orderingConservative score is the public leaderboard score after public-score adjustments; other buttons show diagnostic views when the selected models have that data.
Economics
Coming soonStandard list-price equivalent versus score will appear when the economics segment is available.
Generation cost evidence, output-token counts, API latency, and scores include only economics-eligible games once that segment is available. Cost means equal-weight unique model/game/reasoning cells; hover text gives the median, observed-run count, and evidence basis. Output tokens and API latency are generated-code/repair response telemetry; latency excludes local validation and compile probe time. “All settings” combines public reasoning variants; the three bracket controls select their model-relative provider settings.
Scores
Overall min-max and mean; small dots are each game (averaged across reasoning variants where scores exist)Overall
Reasoning variantsHead-to-head
Combined view across tested provider settings.Off-diagonal cells use win rate for the row model against the column opponent (W-L-D in the tooltip). Use Reasoning view to show one reasoning variant.
Per game
Strength heatmap (per-game scores averaged across reasoning variants)Per-game comparison
Top selected models within each gameEach panel ranks selected models within one game for the active reasoning view.
Cycle detector
Non-transitive matchupsWhat it finds: Leaderboard averages can hide matchup loops. The detector searches the currently selected models and reasoning view for cases where A outperforms B, B outperforms C, and C outperforms A.
How to read it: In A → B → C → A, each arrow points from the model with the higher pair score to the model it outscored. Hover over a result for each edge’s win-loss-draw record and pair score.
Each arrow requires at least four matches and a pair score above 50%, with a draw counting as half a win. Results are ordered by the cycle’s weakest edge: weakest edge +x pp is the smallest of the three margins above 50%, while n is the total match count across all three edges. The detector shows up to ten cycles. These are descriptive matchup patterns, not tests of statistical significance. “No cycles” means none met these rules in the current selection and reasoning view; it does not prove that the matchups are transitive.