- Score · relative 0–100
- 36.7
- Rank
- 30
- Status
- Complete 3/3
- Players generated
- 2026-09-30
- Match evidence updated
- 2026-10-06
- Results published
Scores by reasoning setting
Available settings and grouping
Available settings
Checked on 2026-09-30 · Provider documentation
How settings are grouped
- Baseline: Reasoning disabled — tested and included
- Balanced: High — tested and included
- Intensive: Max — tested and included
Scores by game
What this result shows
Strongest included setting: Max (46.1).
Across included settings, the highest game score is Game of the Amazons (52.1); the lowest is Chess (10.4). These are relative scores within each game, not direct comparisons of difficulty.
Match evidence measures these retained programs. It does not establish how a fresh generation will perform. Independent-generation variation is not estimated in this release.
Test protocol and evidence limits
- Model identifier
- fireworks/ember-1
- Provider route (catalog)
- openrouter
- Prompt identity / profile
- Not retained in this result set / Not retained in this result set
- Prompt version
- Not retained in this result set
- Generation policy
- Generation-quality admission policy (see methodology)
- Requested output-token limits
- 262144 (28 of 28 indexed attempts)
Generation evidence
- Reasoning disabled: 8 contributing player programs across 8 games
- High: 8 contributing player programs across 8 games
- Max: 8 contributing player programs across 8 games
Counts describe contributing programs, not independent replications: repairs and resends are not new independent samples. Independent-generation lineage is not retained; the per-game counts below keep unique code identities separate from admitted generation runs.
| Game | Setting | Unique programs | Admitted generation runs | Program score range |
|---|---|---|---|---|
| Chess | high | 1 | 1 | 11.6–11.6 |
| Chess | max | 1 | 1 | 9.1–9.1 |
| Chess | none | 1 | 1 | 10.5–10.5 |
| Game of the Amazons | high | 1 | 1 | 55.3–55.3 |
| Game of the Amazons | max | 1 | 1 | 63.9–63.9 |
| Game of the Amazons | none | 1 | 1 | 37.1–37.1 |
| Go | high | 1 | 1 | 45.0–45.0 |
| Go | max | 1 | 1 | 62.4–62.4 |
| Go | none | 1 | 1 | 42.3–42.3 |
| Hex | high | 1 | 1 | 44.3–44.3 |
| Hex | max | 1 | 1 | 52.2–52.2 |
| Hex | none | 1 | 1 | 29.2–29.2 |
| International Draughts | high | 1 | 1 | 55.1–55.1 |
| International Draughts | max | 1 | 1 | 50.8–50.8 |
| International Draughts | none | 1 | 1 | 47.1–47.1 |
| Shogi | high | 1 | 1 | 36.3–36.3 |
| Shogi | max | 1 | 1 | 54.1–54.1 |
| Shogi | none | 1 | 1 | 17.7–17.7 |
| Tumbleweed | high | 1 | 1 | 32.7–32.7 |
| Tumbleweed | max | 1 | 1 | 66.4–66.4 |
| Tumbleweed | none | 1 | 1 | 16.9–16.9 |
| Xiangqi | high | 1 | 1 | 23.1–23.1 |
| Xiangqi | max | 1 | 1 | 10.3–10.3 |
| Xiangqi | none | 1 | 1 | 8.6–8.6 |
Program ranges describe the retained competitors on this release scale, not a statistical estimate of a fresh generation.
Failures and budgets
0 model-code generation failures; 0 provider/infrastructure failures excluded from the generation denominator. Runtime player faults are separate: 0 recorded unilateral faults across contributing program rows, excluded from rating evidence.
Repair/resend outcomes and available costs are included. Exact retry ceilings and execution budgets are not retained in this result set; see the methodology for the public policy. Do not infer a measured setting from a current provider catalog.
Player program results
Max: 88% / 12% / 0% · n 8High: 75% / 25% / 0% · n 8Reasoning disabled: 88% / 12% / 0% · n 8 All three benchmark settings testedGeneration cost
$0.17average estimated cost per game and reasoning setting
- Median
- $0.15
- Range
- $0.05–$0.51
- Provider-reported total
- $4.11
- Cost data
- 24 combinations · 24 runs
- Price source
- Provider-reported
- Price list
- boardgame-list-prices-2026-09-08
- Average output
- 9.6k tokens
- Output range
- 2.2k–31k tokens
- Runs with output-token data
- 24
How cost is measured
Comparison cost gives equal weight to each model, game, and reasoning-setting combination. Recorded generation attempts include repairs and resends when cost evidence is available; runtime compute and service operation are excluded. Cost and output-token counts use separate telemetry denominators. A list-price estimate is not actual expenditure; mixed evidence is not a uniform standard-price comparison.
Test protocol
- Combines this model's tested reasoning settings
- Player-program generation; retained protocol details are under “Test protocol and evidence limits” above