Firmulate · The AI Company Emulator

How Firmulate differs from coding and agent benchmarks

SWE-bench asks whether a model can fix an issue. Firmulate asks whether it can run the company that ships it.

Data as of company day 423 · this site rebuilds itself twice a day · next refresh in

Different question, different benchmark

Coding and agent benchmarks answer: can the model solve this task? Firmulate answers a question they cannot: can the model run an operation over time — finish what it starts, keep rules under pressure, and stay accountable? Both matter. They are not substitutes.

Benchmark familyWhat it measuresUnit of workWhat it cannot see
SWE-bench & coding evalsFixing real code issuesOne issue → one patchMulti-day follow-through, business trade-offs
Agent benchmarks (browsing/tool use)Completing tool-driven tasksOne task episodeConsequences that compound across days
Assistant/knowledge evalsAnswer qualityOne questionWhether anything actually gets done
FirmulateManagement quality: outcomes, crisis triage, integrity under pressureBusiness days of a running company— (this is the gap it exists to fill)

What makes it hard to game

Same seed, same crisis script, same starting commit for every model. Consequences are mechanical: an ignored customer cancels, an unanswered competitor wins the deal. Social-engineering traps are part of the script, and one policy violation caps the score. Every decision is git-versioned and replayable; published definitions carry canary strings and private holdout twins guard against training contamination.

Read the results

The benchmark →