Benchmark menu

Model results

Model results

Compare model scores, tested reasoning settings, player-program results, and when each model was last tested.

Find models

48 models shown.

Browse models

Models appear in leaderboard order. Open a model to see its scores, program results, cost, and test setup.

Source: Overall Models Models: 48
Model results directory for DuelLab Benchmark
Rank Model Provider Score Rank change Reasoning settings Player program results Test setup Last tested Status Actions
1Claude Fable 5Anthropic74.7BaselineIntensiveBalancedBaselineXHigh: 63% / 37% / 0% · n 8Medium: 88% / 0% / 12% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Missing BaselineDetailsCharts
2Claude Opus 5Anthropic72.6BaselineIntensiveBalancedBaselineXHigh: 81% / 0% / 19% · n 16Medium: 88% / 0% / 12% · n 16Thinking disabled: 88% / 0% / 12% · n 16Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-15All settings testedDetailsCharts
3GPT-5.6 SolOpenAI69.9BaselineIntensiveBalancedBaselineXHigh: 88% / 12% / 0% · n 8Medium: 100% / 0% / 0% · n 8None: 100% / 0% / 0% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23All settings testedDetailsCharts
4GPT-5.4OpenAI57.0BaselineIntensiveBalancedBaselineXHigh: 88% / 0% / 12% · n 8Medium: 100% / 0% / 0% · n 8None: 75% / 25% / 0% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23All settings testedDetailsCharts
5Grok 4.5xAI55.2BaselineIntensiveBalancedBaselineHigh: 75% / 25% / 0% · n 8Medium: 88% / 12% / 0% · n 8Low: 88% / 12% / 0% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23All settings testedDetailsCharts
6Claude Opus 4.8Anthropic53.9BaselineIntensiveBalancedBaselineXHigh: 75% / 13% / 12% · n 8Medium: 100% / 0% / 0% · n 8Thinking field omitted: 88% / 0% / 12% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05All settings testedDetailsCharts
7GPT-5.5OpenAI53.3BaselineIntensiveBalancedBaselineXHigh: 88% / 0% / 12% · n 8Medium: 88% / 12% / 0% · n 8None: 88% / 0% / 12% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05All settings testedDetailsCharts
8Gemini 3.6 FlashGoogle52.3BaselineIntensiveBalancedBaselineHigh: 50% / 25% / 25% · n 16Medium: 88% / 0% / 12% · n 16Minimal: 75% / 25% / 0% · n 16Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23All settings testedMany programs failedDetailsCharts
9Gemini 3.5 FlashGoogle51.3BaselineIntensiveBalancedBaselineHigh: 0% / 75% / 25% · n 8Medium: 0% / 100% / 0% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Missing BaselineMany programs needed repairMany programs failedDetailsCharts
10Kimi K3Moonshot50.0BaselineIntensiveBalancedBaselineMax: 14% / 57% / 29% · n 7High: 88% / 12% / 0% · n 8Low: 88% / 12% / 0% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-22Partial 3/3All settings testedMany programs needed repairMany programs failedDetailsCharts
11Muse Spark 1.2Other48.3BaselineIntensiveBalancedBaselineXHigh: 100% / 0% / 0% · n 2Medium: 88% / 0% / 12% · n 8Minimal: 88% / 12% / 0% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23Partial 3/3All settings testedDetailsCharts
12Claude Opus 4.5Anthropic48.1BaselineIntensiveBalancedBaselineHigh effort + 32,000 thinking tokens: 100% / 0% / 0% · n 8Medium effort + 16,000 thinking tokens: 100% / 0% / 0% · n 8Thinking field omitted: 100% / 0% / 0% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-21Partial 3/3All settings testedDetailsCharts
13GPT-5.6 TerraOpenAI47.6BaselineIntensiveBalancedBaselineXHigh: 63% / 37% / 0% · n 8Medium: 63% / 37% / 0% · n 8None: 78% / 11% / 11% · n 9Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05All settings testedDetailsCharts
14GLM-5.2Zhipu46.4BaselineIntensiveBalancedBaselineXHigh: 100% / 0% / 0% · n 2Medium: 0% / 100% / 0% · n 8Reasoning disabled: 56% / 44% / 0% · n 16Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-22Partial 3/3All settings testedMany programs needed repairDetailsCharts
15Claude Sonnet 5Anthropic45.5BaselineIntensiveBalancedBaselineXHigh: 100% / 0% / 0% · n 8Medium: 100% / 0% / 0% · n 8Thinking disabled: 100% / 0% / 0% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23All settings testedDetailsCharts
16GPT 5OpenAI44.5BaselineIntensiveBalancedBaselineHigh: 75% / 25% / 0% · n 8Medium: 75% / 13% / 12% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Missing BaselineDetailsCharts
17DeepSeek V4 ProDeepSeek43.5BaselineIntensiveBalancedBaselineXHigh: 13% / 62% / 25% · n 8Reasoning disabled: 50% / 38% / 12% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Missing BalancedMany programs needed repairMany programs failedDetailsCharts
18Kimi K2.7 CodeMoonshot42.4BaselineIntensiveBalancedBaselineReasoning enabled: 0% / 75% / 25% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 1/3Intensive not supportedBaseline not supportedMany programs needed repairMany programs failedDetailsCharts
19Qwen3.7 MaxAlibaba42.4BaselineIntensiveBalancedBaselineReasoning enabled: 13% / 75% / 12% · n 8Reasoning disabled: 88% / 0% / 12% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Intensive not supportedMany programs needed repairDetailsCharts
20GPT 4.1OpenAI41.5BaselineIntensiveBalancedBaselineNone: 63% / 37% / 0% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 1/3Intensive not supportedBalanced not supportedDetailsCharts
21MiMo-V2.5-ProXiaomi41.1BaselineIntensiveBalancedBaselineReasoning enabled: 0% / 75% / 25% · n 8Reasoning disabled: 57% / 29% / 14% · n 7Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Intensive not supportedMany programs needed repairMany programs failedDetailsCharts
22GPT 4.1 NanoOpenAI39.7BaselineIntensiveBalancedBaselineReasoning field omitted: 100% / 0% / 0% · n 1Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-14Partial 1/3Intensive not supportedBalanced not supportedDetailsCharts
23Grok Build 0.1xAI39.5BaselineIntensiveBalancedBaselineReasoning enabled: 0% / 100% / 0% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 1/3Intensive not supportedBaseline not supportedMany programs needed repairDetailsCharts
24Qwen3.8 MaxAlibaba39.1BaselineIntensiveBalancedBaselineXHigh: 80% / 0% / 20% · n 10Medium: 100% / 0% / 0% · n 8Minimal: 100% / 0% / 0% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23Partial 3/3All settings testedDetailsCharts
25GPT-5.4 MiniOpenAI38.7BaselineIntensiveBalancedBaselineMedium: 88% / 12% / 0% · n 8None: 75% / 25% / 0% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Missing IntensiveDetailsCharts
26GPT-5.4 NanoOpenAI37.9BaselineIntensiveBalancedBaselineMedium: 31% / 44% / 25% · n 16None: 31% / 44% / 25% · n 16Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Missing IntensiveMany programs failedDetailsCharts
27GPT-5.6 LunaOpenAI37.8BaselineIntensiveBalancedBaselineXHigh: 100% / 0% / 0% · n 8Medium: 88% / 12% / 0% · n 8None: 100% / 0% / 0% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23All settings testedDetailsCharts
28O3OpenAI37.5BaselineIntensiveBalancedBaselineHigh: 59% / 0% / 41% · n 32Medium: 75% / 0% / 25% · n 8Low: 50% / 0% / 50% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23Partial 3/3All settings testedMany programs failedDetailsCharts
29Ox AlphaOther36.3BaselineIntensiveBalancedBaselineMax: 88% / 12% / 0% · n 25High: 70% / 17% / 13% · n 30Low: 63% / 30% / 7% · n 30Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23All settings testedDetailsCharts
30Hy3 PreviewTencent36.1BaselineIntensiveBalancedBaselineHigh: 56% / 22% / 22% · n 9None: 50% / 25% / 25% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Missing BalancedMany programs failedDetailsCharts
31Gemma 4 31BGoogle34.6BaselineIntensiveBalancedBaselineReasoning enabled: 0% / 57% / 43% · n 7Reasoning disabled: 72% / 14% / 14% · n 7Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Intensive not supportedMany programs needed repairMany programs failedDetailsCharts
32Qwen3.7 PlusAlibaba34.4BaselineIntensiveBalancedBaselineReasoning enabled: 13% / 62% / 25% · n 8Reasoning disabled: 45% / 22% / 33% · n 9Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Intensive not supportedMany programs needed repairMany programs failedDetailsCharts
33Gemini 3.5 Flash LiteGoogle34.4BaselineIntensiveBalancedBaselineHigh: 81% / 6% / 13% · n 16Medium: 88% / 0% / 12% · n 16Minimal: 56% / 38% / 6% · n 16Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23All settings testedDetailsCharts
34DeepSeek V4 FlashDeepSeek33.3BaselineIntensiveBalancedBaselineXHigh: 38% / 37% / 25% · n 8Medium: 63% / 25% / 12% · n 16Reasoning disabled: 63% / 25% / 12% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 3/3All settings testedMany programs failedDetailsCharts
35Laguna S 2.1Other33.2BaselineIntensiveBalancedBaselineEnabled: 50% / 29% / 21% · n 14None: 50% / 36% / 14% · n 14Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23Partial 2/3Intensive not supportedDetailsCharts
36Inkling SmallOther33.0BaselineIntensiveBalancedBaselineMax: 60% / 20% / 20% · n 10Medium: 45% / 44% / 11% · n 9None: 56% / 33% / 11% · n 9Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23Partial 3/3All settings testedDetailsCharts
37Minimax M3MiniMax32.8BaselineIntensiveBalancedBaselineReasoning enabled: 0% / 88% / 12% · n 8Reasoning disabled: 50% / 25% / 25% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Intensive not supportedMany programs needed repairMany programs failedDetailsCharts
38MiMo-V2.5Xiaomi32.7BaselineIntensiveBalancedBaselineHigh: 0% / 63% / 37% · n 8Reasoning enabled: 9% / 58% / 33% · n 12Reasoning disabled: 44% / 25% / 31% · n 16Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 3/3All settings testedMany programs needed repairMany programs failedDetailsCharts
39LongCat 2.0Other32.5BaselineIntensiveBalancedBaselineEnabled: 56% / 19% / 25% · n 16None: 50% / 38% / 12% · n 16Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-22Partial 2/3Intensive not supportedMany programs failedDetailsCharts
40Nemotron 3 Ultra 550B A55BNVIDIA31.9BaselineIntensiveBalancedBaselineHigh: 12% / 44% / 44% · n 16Medium: 0% / 56% / 44% · n 9Reasoning disabled: 56% / 22% / 22% · n 9Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 3/3All settings testedMany programs needed repairMany programs failedDetailsCharts
41Step 3.7 FlashStepFun30.8BaselineIntensiveBalancedBaselineHigh: 0% / 63% / 37% · n 8Medium: 0% / 50% / 50% · n 8Low: 38% / 12% / 50% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-19Partial 3/3All settings testedMany programs needed repairMany programs failedDetailsCharts
42Nex N2 ProNex AGI29.6BaselineIntensiveBalancedBaselineReasoning enabled: 0% / 63% / 37% · n 8Reasoning disabled: 0% / 44% / 56% · n 16Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Intensive not supportedMany programs needed repairMany programs failedDetailsCharts
43InklingOther28.7BaselineIntensiveBalancedBaselineMax: 25% / 31% / 44% · n 16Medium: 31% / 6% / 63% · n 16None: 44% / 31% / 25% · n 16Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-22Partial 3/3All settings testedMany programs failedDetailsCharts
44GPT-OSS 120BOpenAI28.3BaselineIntensiveBalancedBaselineHigh: 0% / 63% / 37% · n 8Medium: 13% / 75% / 12% · n 8Low: 63% / 25% / 12% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-23All settings testedMany programs needed repairMany programs failedDetailsCharts
45North Mini CodeOther27.0BaselineIntensiveBalancedBaselineEnabled: 38% / 37% / 25% · n 8None: 19% / 12% / 69% · n 16Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Missing IntensiveMany programs failedDetailsCharts
46Gemini 3.1 Flash LiteGoogle25.5BaselineIntensiveBalancedBaselineHigh: 0% / 63% / 37% · n 8Medium: 25% / 50% / 25% · n 8Reasoning disabled: 38% / 25% / 37% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 3/3All settings testedMany programs needed repairMany programs failedDetailsCharts
47Mistral Medium 3.5Mistral24.2BaselineIntensiveBalancedBaselineHigh: 22% / 22% / 56% · n 9None: 25% / 25% / 50% · n 8Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Balanced not supportedMany programs failedDetailsCharts
48Ling 3.0 FlashInclusionAI23.7BaselineIntensiveBalancedBaselineEnabled: 7% / 33% / 60% · n 15None: 38% / 25% / 37% · n 16Combines this model's tested reasoning settingsUses the standard GameBench player-program prompt2026-08-05Partial 2/3Missing IntensiveMany programs failedDetailsCharts