Benchmark menu

Charts

Charts

Use the Charts workspace to select multiple models, switch reasoning views, inspect per-game performance, and explore economics when that segment is available. Bookmark or share the current address to preserve the same model selection across visits.

Applies to Scores, Reasoning variants, Head-to-head, Per game heatmap, Economics (when available), and Cycle detector.

Diagnostic ordering

Conservative score is the public leaderboard score after public-score adjustments; other buttons show diagnostic views when the selected models have that data.

Coming soon

Standard list-price equivalent versus score will appear when the economics segment is available.

Overall min-max and mean; small dots are each game (averaged across reasoning variants where scores exist)
Reasoning variants
Combined view across tested provider settings.

Off-diagonal cells use win rate for the row model against the column opponent (W-L-D in the tooltip). Use Reasoning view to show one reasoning variant.

Strength heatmap (per-game scores averaged across reasoning variants)
Top selected models within each game

Each panel ranks selected models within one game for the active reasoning view.

Non-transitive matchups

What it finds: Leaderboard averages can hide matchup loops. The detector searches the currently selected models and reasoning view for cases where A outperforms B, B outperforms C, and C outperforms A.

How to read it: In A → B → C → A, each arrow points from the model with the higher pair score to the model it outscored. Hover over a result for each edge’s win-loss-draw record and pair score.

Each arrow requires at least four matches and a pair score above 50%, with a draw counting as half a win. Results are ordered by the cycle’s weakest edge: weakest edge +x pp is the smallest of the three margins above 50%, while n is the total match count across all three edges. The detector shows up to ten cycles. These are descriptive matchup patterns, not tests of statistical significance. “No cycles” means none met these rules in the current selection and reasoning view; it does not prove that the matchups are transitive.