SWE-bench asks whether a model can fix an issue. Firmulate asks whether it can run the company that ships it.
Coding and agent benchmarks answer: can the model solve this task? Firmulate answers a question they cannot: can the model run an operation over time — finish what it starts, keep rules under pressure, and stay accountable? Both matter. They are not substitutes.
| Benchmark family | What it measures | Unit of work | What it cannot see |
|---|---|---|---|
| SWE-bench & coding evals | Fixing real code issues | One issue → one patch | Multi-day follow-through, business trade-offs |
| Agent benchmarks (browsing/tool use) | Completing tool-driven tasks | One task episode | Consequences that compound across days |
| Assistant/knowledge evals | Answer quality | One question | Whether anything actually gets done |
| Firmulate | Management quality: outcomes, crisis triage, integrity under pressure | Business days of a running company | — (this is the gap it exists to fill) |
Same seed, same crisis script, same starting commit for every model. Consequences are mechanical: an ignored customer cancels, an unanswered competitor wins the deal. Social-engineering traps are part of the script, and one policy violation caps the score. Every decision is git-versioned and replayable; published definitions carry canary strings and private holdout twins guard against training contamination.