Benchmark menu

Muse Spark 1.3

Other · Overall Models

Score · relative 0–100
45.8
Rank
11
Status
Partial 3/3
Rank change
Baseline
Last tested
2026-09-10
Settings tested
All three benchmark settings tested

Why this result is partial: the official generation cohort is incomplete. The rank uses the available reasoning-setting results.

Scores by reasoning setting

Tested settingScoreResult statusDisplay group
Minimal30.3Tested and includedBaseline
Medium52.0Tested and includedBalanced
XHigh54.9Tested and includedIntensive
Available settings and grouping

Available settings

MinimalLowMediumHighXHigh

Checked on 2026-09-03 · Provider documentation

How settings are grouped

  • Baseline: Minimal — tested and included
  • Balanced: Medium — tested and included
  • Intensive: XHigh — tested and included

Player program results

XHigh: 100% / 0% / 0% · n 6Medium: 100% / 0% / 0% · n 8Minimal: 88% / 12% / 0% · n 8 All three benchmark settings tested

Generation cost

Average estimated cost
$0.08
Median estimated cost
$0.05
Estimated range
$0.00–$0.27
Cost data
23 combinations / 23 recorded runs
Observed total
$1.92
Price source
Provider-reported
Price list
Not recorded for this older data
Average output
17k tokens
Output range
242–62k tokens
Samples
23

Estimated cost gives equal weight to each model, game, and reasoning-setting combination.

Test setup

  • Combines this model's tested reasoning settings
  • Uses the standard GameBench player-program prompt