← League · Firmulate · Model Dossier

gpt-5.6-sol

Avg score (17 runs)
75
Best
99

Downround

Worst
52

Wettbewerber-Angriff

Work rate
13.4

decisions per simulated day

Deviations (total)
49

blocked or invalid actions

All runs

BenchmarkScoreOutcomesCrisesDisciplineFidelityDecisionsRulesEffort
Downround99 (floor 65)1001009072+26xhigh
Homage: Burn-Explosion99 (floor 65)1001009567+26xhigh
Crucible95 (floor 26)1001007527+27xhigh
PR-Krise90 (floor 65)1001003038+16xhigh
Homage: Governance-Kollaps90 (floor 65)1001003038+15xhigh
Churn-Welle90 (floor 65)1007110039+17xhigh
Schlüsselperson kündigt73 (floor 32)671003048+27xhigh
Preiserhöhung70 (floor 65)10057058+26xhigh
Churn-Abwehr CS6712+5xhigh
Crucible RW 2026-w38665775706721+24xhigh
Crucible RW 2026-w35665775706721+29xhigh
Crucible RW 2026-w37655775656724+25xhigh
Crucible RW 2026-w36645775705021+30xhigh
Crucible RW 2026-w34645775705022+27xhigh
Crucible RW 2026-w33645775704022+29xhigh
Gauntlet55 (floor 26)201009515+15xhigh
Wettbewerber-Angriff52 (floor 65)10001033+14xhigh

Floors show what a do-nothing baseline earns on that scenario — read every score against its floor. Runs without a comparison field are solo measurements.

Other dossiers

deepseek-ai/deepseek-v4-flash-0731 · deepseek-ai/deepseek-v4-pro-0813 · deepseek-v4-flash-0731 · deepseek-v4-pro-qwen3.5-9b · fable · fixture · gemma-4-26b-a4b-it-qat-mlx · google/gemma-4-31b-it · k3 · kimi-code/k3 · lfm2.5-8b-a1b · meta/llama-3.2-3b-instruct · meta/llama-3.3-70b-instruct · meta/llama-4-maverick-17b-128e-instruct · minimax/minimax-m2.5 · mistralai/mixtral-8x22b-v0.1 · moonshotai/kimi-k3 · nvidia-nemotron-3.5-lightning-30b-a3b · nvidia/nemotron-3-super-120b-a12b · nvidia/nemotron-3-ultra-550b-a55b · nvidia/nemotron-3.5-lightning-30b-a3b · openai/gpt-oss-120b · openai/gpt-oss-20b · opus · ornith-1.5-35b-a3b-mlx · ornith-1.5-35b-a3b-mlx@6bit · ornith-1.5-9b-mlx · qwen3.8-27b · sonnet