Benchmark menu

GPT 4.1 Nano

OpenAI · Overall Models

Score · relative 0–100
40.1
Rank
28
Status
Partial 1/3
Rank change
Baseline
Players generated
2026-07-12 – 2026-07-12
Match evidence updated
2026-09-26
Results published
2026-09-26T08:55:16Z
Settings tested
Intensive not supported; Balanced not supported

Why this result is partial: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing; the official generation cohort is incomplete. The rank uses the available reasoning-setting results.

What this result shows

Strongest included setting: Reasoning field omitted (40.1).

Per-game evidence is unavailable.

Match evidence measures these retained programs. It does not establish how a fresh generation will perform. Independent-generation variation is not estimated in this release.

Test protocol and evidence limits
Model identifier
gpt-4.1-nano
Provider route (catalog)
openai
Prompt identity / profile
Not retained in this result set / Not retained in this result set
Prompt version
Not retained in this result set
Generation policy
Generation-quality admission policy (see methodology)
Requested output-token limits
Not retained (0 of 2 indexed attempts)

Generation evidence

  • Reasoning field omitted: 1 contributing player programs across 1 games

Counts describe contributing programs, not independent replications: repairs and resends are not new independent samples. Independent-generation lineage is not retained; the per-game counts below keep unique code identities separate from admitted generation runs.

GameSettingUnique programsAdmitted generation runsProgram score range
Gonone1140.1–40.1

Program ranges describe the retained competitors on this release scale, not a statistical estimate of a fresh generation.

Failures and budgets

0 model-code generation failures; 0 provider/infrastructure failures excluded from the generation denominator. Runtime player faults are separate: 0 recorded unilateral faults across contributing program rows, excluded from rating evidence.

Repair/resend outcomes and available costs are included. Exact retry ceilings and execution budgets are not retained in this result set; see the methodology for the public policy. Do not infer a measured setting from a current provider catalog.

Scores by reasoning setting

Tested settingScoreResult statusDisplay group
Reasoning field omitted40.1Tested and includedBaseline
Available settings and grouping

Available settings

Reasoning field omitted

Checked on 2026-07-31 · Provider documentation

How settings are grouped

  • Baseline: Reasoning field omitted — tested and included
  • Balanced: Not available for this model
  • Intensive: Not available for this model

Player program results

Reasoning field omitted: 100% / 0% / 0% · n 1 Intensive not supportedBalanced not supported

Generation cost

Average estimated cost
$0.0023
Median estimated cost
$0.0023
Estimated range
$0.0023–$0.0023
Cost data
1 combinations / 1 recorded runs
Known cost subtotal (incomplete or unverified charges)
$0.0023
Price source
Combined: Recorded attempt + Standard list-price estimate
Price list
boardgame-list-prices-2026-07-12 + boardgame-list-prices-2026-09-08
Average output
2.0k tokens
Output range
2.0k–2.0k tokens
Runs with output-token evidence
1

Comparison cost gives equal weight to each model, game, and reasoning-setting combination. Recorded generation attempts include repairs and resends when cost evidence is available; runtime compute and service operation are excluded. Cost and output-token counts use separate telemetry denominators. A list-price estimate is not actual expenditure; mixed evidence is not a uniform standard-price comparison.

Test protocol

  • Combines this model's tested reasoning settings
  • Player-program generation; retained protocol evidence below