Benchmark menu

Mercury 2.5 Preview

Other · Overall Models

Score · relative 0–100
18.4
Rank
62
Status
Complete 3/3
Rank change
Baseline
Last tested
2026-09-10
Settings tested
All three benchmark settings tested

Scores by reasoning setting

Tested settingScoreResult statusDisplay group
None21.3Tested and includedBaseline
Medium17.7Tested and includedBalanced
High16.3Tested and includedIntensive
Available settings and grouping

Available settings

NoneLowMediumHigh

Checked on 2026-09-03 · Provider documentation

How settings are grouped

  • Baseline: None — tested and included
  • Balanced: Medium — tested and included
  • Intensive: High — tested and included

Player program results

High: 63% / 12% / 25% · n 8Medium: 75% / 0% / 25% · n 8None: 25% / 38% / 37% · n 8 All three benchmark settings testedMany programs failed

Generation cost

Average estimated cost
$0.0012
Median estimated cost
$0.000945
Estimated range
$0.000440–$0.0026
Cost data
24 combinations / 24 recorded runs
Observed total
$0.03
Price source
Provider-reported
Price list
Not recorded for this older data
Average output
4.9k tokens
Output range
1.0k–11k tokens
Samples
24

Estimated cost gives equal weight to each model, game, and reasoning-setting combination.

Test setup

  • Combines this model's tested reasoning settings
  • Uses the standard GameBench player-program prompt