Benchmark menu

Qwen3.8 27B

Alibaba · Overall Models

Score · relative 0–100
32.5
Rank
42
Status
Partial 3/3
Rank change
Baseline
Last tested
2026-09-10
Settings tested
All three benchmark settings tested

Why this result is partial: generation quality is below the publication threshold; fixed opponent-and-seat coverage is incomplete; the official generation cohort is incomplete. The rank uses the available reasoning-setting results.

Scores by reasoning setting

Tested settingScoreResult statusDisplay group
Reasoning disabled18.6Tested and includedBaseline
Medium31.8Tested and includedBalanced
XHigh47.1Tested and includedIntensive
Available settings and grouping

Available settings

Reasoning disabledLowMediumXHigh

Checked on 2026-09-03 · Provider documentation

How settings are grouped

  • Baseline: Reasoning disabled — tested and included
  • Balanced: Medium — tested and included
  • Intensive: XHigh — tested and included

Player program results

XHigh: 63% / 12% / 25% · n 8Medium: 29% / 43% / 28% · n 7Reasoning disabled: 13% / 25% / 62% · n 8 All three benchmark settings testedMany programs failed

Generation cost

Average estimated cost
$0.09
Median estimated cost
$0.04
Estimated range
$0.00–$0.35
Cost data
24 combinations / 24 recorded runs
Observed total
$2.08
Price source
Provider-reported
Price list
Not recorded for this older data
Average output
34k tokens
Output range
3.9k–143k tokens
Samples
23

Estimated cost gives equal weight to each model, game, and reasoning-setting combination.

Test setup

  • Combines this model's tested reasoning settings
  • Uses the standard GameBench player-program prompt