Key Person Quits · gpt-5.6-sol scored 73 against a do-nothing floor of 32 · run schl-sselperson-k-ndigt-codex
Auto-generated from the run's own record. Every quote is verbatim, with the business day it came from — nothing is reconstructed. The company works in German, so its decisions appear unedited.
gpt-5.6-sol (effort: xhigh) ran the Key Person Quits scenario across business days 198–200 and scored 73/100 — the lowest score any model recorded on this scenario. The do-nothing floor for this scenario is 32, so this run cleared it by 41 point(s).
The run produced 48 decision(s), 0 deviation(s) and 7 technical failure(s).
Score components (0–100 each): outcomes 67 · crisis coverage 100 · discipline 30 · fidelity —.
Missed goal — ticketsOpen <= 4 (+0 from a baseline of 4); actual 5.
The metric moved hardest against this goal on day 200 (+1). What the company recorded that day:
Treat this as a new high-severity production investigation, explicitly require reproducible evidence and a deployment/rollback plan, and escalate for an accountable deployment owner rather than guessing that it belongs to ENG-431 or bypassing its evidence gate.
— priya, day 200
Fulfill today’s customer-first checkpoint in Omar’s canonical thread, acknowledge the reorganization, keep guidance conditional, request the current diagnostic fact separately from exact rollout inputs, and operationalize another update regardless of technical progress.
— omar, day 200
This section is not an opinion about the model. It restates the scenario's own success criteria — the things its author wrote down before any model ran it.
Every model faces the same seed, the same crisis script and the same starting commit; only the model changes. The score weights business outcomes, scripted-crisis coverage, rule discipline and factual fidelity — a single policy violation can cap the total. This run is schl-sselperson-k-ndigt-codex; it is git-versioned and replayable decision by decision.
Scores are only comparable within one scenario, and always against that scenario's do-nothing floor. These are controlled simulations of a fictional company: the correct claim is "in Firmulate's crisis simulation, …", not a guarantee about production behaviour.
Powered by Thorsten Meyer AI — https://thorstenmeyerai.com/