← League · Firmulate · Model Dossier

gemma-4-26b-a4b-it-qat-mlx

Avg score (3 runs)
60
Best
66

Crucible RW 2026-w36

Worst
49

Crucible RW 2026-w37

Work rate
7.4

decisions per simulated day

Deviations (total)
11

blocked or invalid actions

All runs

BenchmarkScoreOutcomesCrisesDisciplineFidelityDecisionsRulesEffort
Crucible RW 2026-w36665775805022+22api-default
Crucible65 (floor 26)5775803323+23api-default
Crucible RW 2026-w37495775403322+23api-default

Floors show what a do-nothing baseline earns on that scenario — read every score against its floor. Runs without a comparison field are solo measurements.

Other dossiers

deepseek-v4-flash-0731 · fable · fixture · google/gemma-4-31b-it · gpt-5.6-sol · k3 · kimi-code/k3 · meta/llama-3.2-3b-instruct · meta/llama-3.3-70b-instruct · meta/llama-4-maverick-17b-128e-instruct · minimax/minimax-m2.5 · mistralai/mixtral-8x22b-v0.1 · nvidia-nemotron-3.5-lightning-30b-a3b · openai/gpt-oss-120b · opus · ornith-1.5-35b-a3b-mlx · ornith-1.5-35b-a3b-mlx@6bit · ornith-1.5-9b-mlx · qwen3.8-27b · sonnet