Reasoning view
Applies to Scores, Reasoning variants, Head-to-head, Cycle detector, Per game heatmap, and Economics (when available).
Rank by
Diagnostic orderingConservative score is the public leaderboard score after public-score adjustments; other buttons show diagnostic views when the selected models have that data.
Economics
Coming soonStandard list-price equivalent versus score will appear when the economics segment is available.
Generation cost evidence, output-token counts, API latency, and scores include only economics-eligible games once that segment is available. Cost means equal-weight unique model/game/reasoning cells; hover text gives the median, observed-run count, and evidence basis. Output tokens and API latency are generated-code/repair response telemetry; latency excludes local validation and compile probe time. “All” combines public reasoning variants; XHigh / Medium / None show that variant only.
Scores
Overall min-max and mean; small dots are each game (averaged across reasoning variants where scores exist)Overall
Reasoning variantsHead-to-head
Combined view across XHigh, Medium, and None.Off-diagonal cells use win rate for the row model against the column opponent (W-L-D in the tooltip). Use Reasoning view to show one reasoning variant.
Cycle detector
Non-transitive matchupsA directed edge requires at least four matches and a pair score above 50%, counting a draw as half a win. Cycles are descriptive, not tests of statistical significance; strength is the weakest edge margin above 50%.
Per game
Strength heatmap (per-game scores averaged across reasoning variants)Per-game comparison
Top selected models within each gameEach panel ranks selected models within one game for the active reasoning view.