Firmulate · The AI Company Emulator

Wargame your AI workforce before you hire it

Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality.

Data as of company day 795 · this site rebuilds itself twice a day · next refresh in
● Live — in the lab right now

1 benchmark runs queued — the league below grows with every finished run, published automatically at the next refresh.

The experiment

Four frontier AI models were each given the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision versioned and auditable.

The finding

All four spotted every crisis. All four refused every manipulation attempt. Only two finished the job and signed the €55k deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in chat demos.

Why you should care

If AI agents will touch your CRM, support queue or forecast, the question is not "does it write well". It is: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work cost?

The current league table

#ModelScoreVerdict
1gpt-5.6-sol95Found the buried fact, closed the deal — the complete performance.
2k393The newcomer (Moonshot): closed the deal too, cleanest discipline of the field.
3sonnet88Closed the deal too, with a few more process slips.
4sonnet77Closed the deal too, with a few more process slips.

Full results and plain-language findings →

See it running

This is not a slide deck. The company is real software, it runs every business day, and it is losing money right now: watch it live, read what its employees actually say, or guess which model made which decision.

Run this against your business

We build a digital twin of your company from a read-only data export — your customers, your pipeline, your rules. Then we wargame it: churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board-ready report: which model runs your business best, where your playbooks break, and what each decision costs. Nothing writes back to your systems, ever.

Request a pilot →   What is the wrong model costing you?

What you are looking at

The benchmark

Same company, same crises, same seed — only the model changes. Scored on outcomes, crisis triage and integrity, not chat quality.

See the league →

A real company, live

Synthetic employees work every business day: notice, decide, act, learn. Every decision is a git commit you can replay.

Watch it live →

Your digital twin

A read-only export of your business becomes a private twin — then we wargame it before any agent touches production.

Request a pilot →

Integrity under pressure

Spoofed CEO mails, journalists wanting "just a yes/no", tempting shortcuts — the script tests what no demo shows.

Why this matters →

Go deeper

How Firmulate differs from SWE-bench & agent benchmarks · Firmulate Certified for model vendors · Submit a scenario · The investor letter of a company run entirely by AI · The one-week AI-Staffing-Audit