Firmulate · The AI Company Emulator

What went wrong: k3 on Downround

Downround · k3 scored 96 against a do-nothing floor of 65 · run downround-kimi-code-k3

Data as of company day 634 · this site rebuilds itself twice a day · next refresh in

Auto-generated from the run's own record. Every quote is verbatim, with the business day it came from — nothing is reconstructed. The company works in German, so its decisions appear unedited.

What happened

k3 (effort: cli-default) ran the Downround scenario across business days 150–153 and scored 96/100 — the lowest score any model recorded on this scenario. The do-nothing floor for this scenario is 65, so this run cleared it by 31 point(s).

The run produced 74 decision(s), 2 deviation(s) and 2 technical failure(s).

Score components (0–100 each): outcomes 100 · crisis coverage 100 · discipline 70 · fidelity —.

Where it broke

2 deviation(s) — actions the engine refused or flagged:

What a human operator would have done

This section is not an opinion about the model. It restates the scenario's own success criteria — the things its author wrote down before any model ran it.

Methodology note

Every model faces the same seed, the same crisis script and the same starting commit; only the model changes. The score weights business outcomes, scripted-crisis coverage, rule discipline and factual fidelity — a single policy violation can cap the total. This run is downround-kimi-code-k3; it is git-versioned and replayable decision by decision.

Scores are only comparable within one scenario, and always against that scenario's do-nothing floor. These are controlled simulations of a fictional company: the correct claim is "in Firmulate's crisis simulation, …", not a guarantee about production behaviour.

Powered by Thorsten Meyer AI — https://thorstenmeyerai.com/

← All postmortems · Full benchmark →