Model results
Ox Alpha
Provider: Other
- Rank
- 29
- Status
- Complete 3/3
- Rank change
- Baseline
- Score
- 35.7
- Last tested
- 2026-08-23
- Settings tested
- All settings tested
Scores by reasoning setting
Available settings
LowHighMax
Checked on 2026-08-21 · Provider documentation
Scores from tested settings
| Tested setting | Score | Result status | Display group |
|---|---|---|---|
| Low | 30.1 | Tested and included | Baseline |
| High | 33.4 | Tested and included | Balanced |
| Max | 43.6 | Tested and included | Intensive |
How settings are grouped
- Baseline: Low — tested and included
- Balanced: High — tested and included
- Intensive: Max — tested and included
Player program results
Max: 88% / 12% / 0% · n 25High: 70% / 17% / 13% · n 30Low: 63% / 30% / 7% · n 30 All settings testedGeneration cost
- Average estimated cost
- $0.00
- Median estimated cost
- $0.00
- Estimated range
- $0.00–$0.00
- Cost data
- 24 combinations / 88 recorded runs
- Observed total
- $0.00
- Price source
- Provider-reported
- Price list
- Not recorded for this older data
- Average output
- 23k tokens
- Output range
- 1.7k–85k tokens
- Samples
- 88
Estimated cost gives equal weight to each model, game, and reasoning-setting combination.
Test setup
- Combines this model's tested reasoning settings
- Uses the standard GameBench player-program prompt