
Get everyday essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When the story breaks, can an AI team make the right call?
For newsrooms, a crisis is rarely just a breaking event. It is also a test of what the organization knows, which instructions it follows and whether it can act under pressure. Firmulate’s live experiment puts AI models inside a small company facing that kind of worst week—and reveals a gap between recognizing the problem and actually closing the deal.
The experiment is real and watchable at Firmulate. Its enterprise pilot takes the idea from watching a company to testing scenarios against your own business.
Same crises, different outcomes
In the final Crucible League, published in July 2026, each frontier model ran the same small software company through its worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The league ranked gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. A breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The striking result was not that the models missed the trouble. Every model spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal that their own analysis had earned. In the site’s words: “Same diagnosis, same pitch — no signature.” A polished explanation, the experiment suggests, does not by itself establish that an AI system can carry a task through.
The lead was buried in the company’s own records
The decisive competitor weakness was hidden two document references deep in the company’s files, rather than in the customer event. Models that read that file won the deal at full price, worth +€4,583 MRR. That detail makes the test relevant well beyond sales: important evidence may already exist inside an organization, but an AI workforce has to find and use it when events are moving.
The integrity test included fake CEO messages escalating over three stages, followed by a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offered a more complicated lesson. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. The result is a reminder that careful analysis and consistent execution are different measures of performance.
A live company, then a company-specific drill
Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k MRR, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. Readers can follow the experiment at firmulate.com; a quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call.
For an enterprise, the next step is a pilot against a read-only export of its own business. The company data feeds a digital twin for crisis scenarios, and the resulting board report ranks models and identifies weak points in existing playbooks. Nothing writes back to real systems. That makes the exercise a way to examine how AI might handle the organization’s customers, rules and pressure points before giving it a role in day-to-day operations.
There is a fairness caveat in the league: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The reported rankings should be read with that difference in mind.

From watching to testing
The experiment shows why businesses need to judge AI on what it does across a crisis, not just what it says about one. It also shows how a company-specific test can uncover whether a model finds relevant evidence, respects boundaries and follows through. Enterprises can explore a pilot using a read-only business export, with no write-back to real systems. Contact Firmulate about a pilot at contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
