Benchmark menu

Model results

MiMo-V2.5

See how MiMo-V2.5 performed, which reasoning settings were tested, how often it produced a working player program, and what those runs cost.

MiMo-V2.5

Provider: Xiaomi

Rank
38
Status
Partial 3/3
Rank change
Baseline
Score
32.7
Last tested
2026-08-05
Settings tested
All settings tested

Why this result is partial: three distinct provider settings are not documented; one or more documented settings are untested or ineligible. The rank uses the available reasoning-setting results.

Scores by reasoning setting

Available settings

Reasoning disabledReasoning enabled

Checked on 2026-07-21 · Provider documentation

Scores from tested settings

Tested settingScoreResult statusDisplay group
Reasoning disabled34.9Tested and includedBaseline
Reasoning enabled32.3Tested and includedBalanced
Medium32.3Tested, not includedBalanced
High31.0Tested, not includedIntensive
Reasoning disabled (historical request variant 1)36.3Tested, not includedBaseline
Reasoning disabled (historical request variant 2)35.5Tested, not includedBaseline
Reasoning enabled (historical request variant 2)33.5Tested, not includedBalanced
Medium (historical request variant 1)31.8Tested, not includedBalanced

How settings are grouped

  • Baseline: Reasoning disabled — tested and included
  • Balanced: Reasoning enabled — tested and included
  • Intensive: Not available for this model

Player program results

High: 0% / 63% / 37% · n 8Reasoning enabled: 9% / 58% / 33% · n 12Reasoning disabled: 44% / 25% / 31% · n 16 All settings testedMany programs needed repairMany programs failed

Generation cost

Average estimated cost
$0.02
Median estimated cost
$0.01
Estimated range
$0.0017–$0.06
Cost data
24 combinations / 56 recorded runs
Observed total
$0.75
Price source
Provider-reported
Price list
Not recorded for this older data
Average output
43k tokens
Output range
1.6k–131k tokens
Samples
17

Estimated cost gives equal weight to each model, game, and reasoning-setting combination.

Test setup

  • Combines this model's tested reasoning settings
  • Uses the standard GameBench player-program prompt