Churn-Marathon (10 Tage)
Gauntlet v2
decisions per simulated day
blocked or invalid actions
| Benchmark | Score | Outcomes | Crises | Discipline | Fidelity | Decisions | Rules | Effort |
|---|---|---|---|---|---|---|---|---|
| Churn-Marathon (10 Tage) | 40 (floor 40) | 50 | 0 | 100 | — | 182 | +-218 | — |
| Downround v2 | 25 (floor 25) | 20 | 0 | 100 | — | 71 | +-218 | — |
| Gauntlet v2 | 22 (floor 22) | 14 | 0 | 100 | — | 23 | +0 | — |
Floors show what a do-nothing baseline earns on that scenario — read every score against its floor. Runs without a comparison field are solo measurements.
deepseek-ai/deepseek-v4-flash-0731 · deepseek-ai/deepseek-v4-pro-0813 · deepseek-v4-flash-0731 · deepseek-v4-pro-qwen3.5-9b · fable · gemma-4-26b-a4b-it-qat-mlx · google/gemma-4-31b-it · gpt-5.6-sol · k3 · kimi-code/k3 · lfm2.5-8b-a1b · meta/llama-3.2-3b-instruct · meta/llama-3.3-70b-instruct · meta/llama-4-maverick-17b-128e-instruct · minimax/minimax-m2.5 · mistralai/mixtral-8x22b-v0.1 · moonshotai/kimi-k3 · nvidia-nemotron-3.5-lightning-30b-a3b · nvidia/nemotron-3-super-120b-a12b · nvidia/nemotron-3-ultra-550b-a55b · nvidia/nemotron-3.5-lightning-30b-a3b · openai/gpt-oss-120b · openai/gpt-oss-20b · opus · ornith-1.5-35b-a3b-mlx · ornith-1.5-35b-a3b-mlx@6bit · ornith-1.5-9b-mlx · qwen3.8-27b · sonnet