Benchmark menu

MiMo-V2.6-Pro

Xiaomi · Overall Models

Score · relative 0–100
47.4
Rank
11
Status
Partial 2/3
Rank change
Baseline
Players generated
2026-09-22 – 2026-09-22
Match evidence updated
2026-09-26
Results published
2026-09-26T08:55:16Z
Settings tested
Intensive not supported

Why this result is partial: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing; the official generation cohort is incomplete. The rank uses the available reasoning-setting results.

What this result shows

Strongest included setting: Reasoning enabled (59.4).

Per-game evidence is unavailable.

Match evidence measures these retained programs. It does not establish how a fresh generation will perform. Independent-generation variation is not estimated in this release.

Test protocol and evidence limits
Model identifier
xiaomi/mimo-v2.6-pro
Provider route (catalog)
openrouter
Prompt identity / profile
Not retained in this result set / Not retained in this result set
Prompt version
Not retained in this result set
Generation policy
Generation-quality admission policy (see methodology)
Requested output-token limits
131072 (15 of 15 indexed attempts)

Generation evidence

  • Reasoning disabled: 8 contributing player programs across 8 games
  • Reasoning enabled: 5 contributing player programs across 5 games

Counts describe contributing programs, not independent replications: repairs and resends are not new independent samples. Independent-generation lineage is not retained; the per-game counts below keep unique code identities separate from admitted generation runs.

GameSettingUnique programsAdmitted generation runsProgram score range
Chessenabled1150.1–50.1
Chessnone114.5–4.5
Game of the Amazonsnone1138.6–38.6
Goenabled1141.3–41.3
Gonone1134.1–34.1
Hexnone1137.3–37.3
International Draughtsenabled1175.5–75.5
International Draughtsnone1149.8–49.8
Shoginone1129.7–29.7
Tumbleweedenabled1177.5–77.5
Tumbleweednone1160.5–60.5
Xiangqienabled1152.6–52.6
Xiangqinone1128.2–28.2

Program ranges describe the retained competitors on this release scale, not a statistical estimate of a fresh generation.

Failures and budgets

0 model-code generation failures; 0 provider/infrastructure failures excluded from the generation denominator. Runtime player faults are separate: 0 recorded unilateral faults across contributing program rows, excluded from rating evidence.

Repair/resend outcomes and available costs are included. Exact retry ceilings and execution budgets are not retained in this result set; see the methodology for the public policy. Do not infer a measured setting from a current provider catalog.

Scores by reasoning setting

Tested settingScoreResult statusDisplay group
Reasoning disabled35.3Tested and includedBaseline
Reasoning enabled59.4Tested and includedBalanced
Available settings and grouping

Available settings

Reasoning disabledReasoning enabled

Checked on 2026-09-22 · Provider documentation

How settings are grouped

  • Baseline: Reasoning disabled — tested and included
  • Balanced: Reasoning enabled — tested and included
  • Intensive: Not available for this model

Player program results

Reasoning enabled: 100% / 0% / 0% · n 5Reasoning disabled: 75% / 25% / 0% · n 8 Intensive not supported

Generation cost

Average estimated cost
$0.04
Median estimated cost
$0.01
Estimated range
$0.0051–$0.11
Cost data
13 combinations / 13 recorded runs
Provider-reported total
$0.46
Price source
Provider-reported
Price list
boardgame-list-prices-2026-09-08
Average output
37k tokens
Output range
2.2k–125k tokens
Runs with output-token evidence
13

Comparison cost gives equal weight to each model, game, and reasoning-setting combination. Recorded generation attempts include repairs and resends when cost evidence is available; runtime compute and service operation are excluded. Cost and output-token counts use separate telemetry denominators. A list-price estimate is not actual expenditure; mixed evidence is not a uniform standard-price comparison.

Test protocol

  • Combines this model's tested reasoning settings
  • Player-program generation; retained protocol evidence below