Model detail
Claude Opus 5
Provider: Anthropic · Public slug: claude-opus-5
- Rank
- 1
- Rank movement
- Baseline
- Score
- 63.6
- Last measured
- 2026-07-26
- Coverage
- Complete coverage
Reasoning profile
Available reasoning settings
Thinking disabledLowMediumHighXHighMax
Checked on 2026-07-25 · provider source
Tested reasoning settings
| Provider setting | Score | Evidence | Filter bracket |
|---|---|---|---|
| Thinking disabled | 55.4 | Retained request evidence; eligible | Baseline |
| Medium | 70.5 | Retained request evidence; eligible | Balanced |
| XHigh | 64.9 | Retained request evidence; eligible | Intensive |
Filter assignment status
- Baseline: Thinking disabled — tested and eligible
- Balanced: Medium — tested and eligible
- Intensive: XHigh — tested and eligible
Generation health
XHigh: 81% / 0% / 19% · n 16Medium: 88% / 0% / 12% · n 16Thinking disabled: 81% / 0% / 19% · n 16 Complete coverageCost summary
- Mean standard list-price equivalent
- $0.70
- Median unique-cell cost
- $0.48
- Unique-cell range
- $0.22–$2.10
- Coverage
- 24 unique cells / 48 observed runs
- Observed-run total
- $33.78
- Evidence basis
- Standard list-price estimate
- Mean output
- 26k tok
- Output range
- 5.4k–88k tok
- Token samples
- 48
Cost means equal-weight unique model/game/reasoning cells.
Run setup
- Overall Models: combines public reasoning variants
- Standard GameBench code-generation prompt