Score · relative 0–100
49.4
Rank
11
Status
Complete 3/3
Players generated
2026-09-30
Match evidence updated
2026-10-06
Results published

Scores by reasoning setting

  • XHigh (Intensive group, tested and included)83.6
  • Medium (Balanced group, tested and included)36.4
  • Low (Baseline group, tested and included)28.1

Every tested setting is included in the overall score.

Available settings and grouping

Available settings

LowMediumHighXHighMax

Checked on 2026-09-30 · Provider documentation

How settings are grouped

  • Baseline: Low — tested and included
  • Balanced: Medium — tested and included
  • Intensive: XHigh — tested and included

Scores by game

  • Go47.1
  • Chess39.0
  • Shogi44.4
  • Xiangqi28.8
  • Hex46.1
  • International Draughts60.1
  • Game of the Amazons63.8
  • Tumbleweed65.6

Relative scores within each game, across the included settings; they do not compare game difficulty.

What this result shows

Strongest included setting: XHigh (83.6).

Across included settings, the highest game score is Tumbleweed (65.6); the lowest is Xiangqi (28.8). These are relative scores within each game, not direct comparisons of difficulty.

Match evidence measures these retained programs. It does not establish how a fresh generation will perform. Independent-generation variation is not estimated in this release.

Test protocol and evidence limits
Model identifier
claude-sonnet-5-5
Provider route (catalog)
anthropic
Prompt identity / profile
Not retained in this result set / Not retained in this result set
Prompt version
Not retained in this result set
Generation policy
Generation-quality admission policy (see methodology)
Requested output-token limits
128000 (28 of 28 indexed attempts)

Generation evidence

  • Low: 8 contributing player programs across 8 games
  • Medium: 8 contributing player programs across 8 games
  • XHigh: 8 contributing player programs across 8 games

Counts describe contributing programs, not independent replications: repairs and resends are not new independent samples. Independent-generation lineage is not retained; the per-game counts below keep unique code identities separate from admitted generation runs.

GameSettingUnique programsAdmitted generation runsProgram score range
Chesslow1112.9–12.9
Chessmedium117.4–7.4
Chessxhigh1196.6–96.6
Game of the Amazonslow1144.6–44.6
Game of the Amazonsmedium1161.7–61.7
Game of the Amazonsxhigh1185.2–85.2
Golow1138.5–38.5
Gomedium1138.5–38.5
Goxhigh1164.4–64.4
Hexlow1121.0–21.0
Hexmedium1126.9–26.9
Hexxhigh1190.4–90.4
International Draughtslow1128.8–28.8
International Draughtsmedium1170.4–70.4
International Draughtsxhigh1181.2–81.2
Shogilow1128.6–28.6
Shogimedium1115.9–15.9
Shogixhigh1188.6–88.6
Tumbleweedlow1135.6–35.6
Tumbleweedmedium1161.3–61.3
Tumbleweedxhigh11100.0–100.0
Xiangqilow1115.0–15.0
Xiangqimedium119.1–9.1
Xiangqixhigh1162.3–62.3

Program ranges describe the retained competitors on this release scale, not a statistical estimate of a fresh generation.

Failures and budgets

0 model-code generation failures; 0 provider/infrastructure failures excluded from the generation denominator. Runtime player faults are separate: 0 recorded unilateral faults across contributing program rows, excluded from rating evidence.

Repair/resend outcomes and available costs are included. Exact retry ceilings and execution budgets are not retained in this result set; see the methodology for the public policy. Do not infer a measured setting from a current provider catalog.

Player program results

XHigh: 88% / 12% / 0% · n 8Medium: 63% / 37% / 0% · n 8Low: 100% / 0% / 0% · n 8 All three benchmark settings tested

Generation cost

$0.22average estimated cost per game and reasoning setting

Median
$0.06
Range
$0.04–$0.73
Estimated total for recorded runs
$5.20
Cost data
24 combinations · 24 runs
Price source
Standard list-price estimate
Price list
boardgame-list-prices-2026-09-08
Average output
19k tokens
Output range
1.2k–70k tokens
Runs with output-token data
24
How cost is measured

Comparison cost gives equal weight to each model, game, and reasoning-setting combination. Recorded generation attempts include repairs and resends when cost evidence is available; runtime compute and service operation are excluded. Cost and output-token counts use separate telemetry denominators. A list-price estimate is not actual expenditure; mixed evidence is not a uniform standard-price comparison.

Test protocol

  • Combines this model's tested reasoning settings
  • Player-program generation; retained protocol details are under “Test protocol and evidence limits” above