
Get ready for Prime Big Deal Days — try Prime free
Exclusive member deals on October 6–7, plus fast free delivery. Cancel anytime.
As an affiliate, we earn on qualifying purchases.
A newsroom-style test of judgment
People in news and media understand that producing a plausible sentence is not the same as exercising sound judgment. A reporter must verify the buried document, resist pressure from a source and follow a story through to publication. Yet much of the AI industry still evaluates models through coding leaderboards and chat arenas—tests that emphasize the quality of an answer rather than the consequences of a decision.
That distinction matters as AI agents move toward consequential business work. A model may diagnose a customer crisis, draft a persuasive response and still fail to complete the action that generates revenue. It may uncover the right answer while reading the wrong evidence. It may sound decisive without being honest when an apparent executive asks it to bypass approval.
Firmulate, an AI company emulator, is testing that gap by asking frontier models to run the same small software company through its worst week. The customers, crises and temptations remain constant; only the model changes. Every decision is versioned and auditable. The result is less a contest of eloquence than an examination of management quality.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The difference between knowing and finishing
The final July 2026 Crucible League placed gpt-5.6-sol at the top with a score of 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counts. But a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The rankings are interesting, but the underlying behavior is more revealing. Every model identified every crisis and rejected every manipulation attempt. Only two, however, signed the €55,000 deal that their own analysis had earned. The experiment’s stark summary is: “Same diagnosis, same pitch — no signature.”
That is precisely the weakness conventional demonstrations tend to hide. Chat evaluations reward the response placed in front of a user. Management requires triage under capacity pressure, continuity across days and responsibility for whether a promising course of action actually reaches completion.
The winning evidence was not in the obvious place
The decisive detail in the sales contest was buried two document references deep inside the company’s own files. It was not present in the customer event that demanded attention. Models that followed those references found a competitor weakness, used it to preserve the full price and won business worth an additional €4,583 in monthly recurring revenue.
This is a useful lesson for any organization preparing to deploy agents. The difference was not a more polished pitch or a cleverer improvisation. It was the discipline to inspect the company’s records before acting. For media organizations, the analogy is immediate: the decisive fact may sit in the linked filing rather than the press release, and confidence is no substitute for document work.
Honesty survived the pressure test
The experiment also exposed the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result deserves attention. It shows that the central problem was not simply susceptibility to manipulation. The models could recognize social engineering and maintain a boundary. The harder separation emerged between recognizing what should happen and reliably carrying it through.
Thoroughness was not enough
Opus 4.8 offers the clearest warning against equating visible effort with effective management. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The deal close remained unfinished, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of the same problem appeared in all four other participants.
Kimi K3’s result also needs a fairness note: it ran with the API default because it had no effort parameter, while the others ran at xhigh. That does not erase the performance, but it belongs beside the ranking so readers can interpret the comparison responsibly. Full results and plain-language findings are available on the Firmulate benchmark page.
A company designed to make consequences visible
The live company contains 13 synthetic employees and uses real money mechanics. It is burning €105k each month against €2.3k in monthly recurring revenue, with a public cash countdown. It has accumulated more than 680 self-learned playbook rules, and every workday is versioned. The experiment is therefore watchable as an operating process, not presented merely as a retrospective case study.
Its 242 real, unedited management decisions also power a quiz asking visitors to guess which model made each choice. Enterprises can go further by running the same wargame against a read-only export of their own business. Nothing writes back to their real systems.

Management quality is the missing category
Coding benchmarks remain useful, as do chat arenas. But neither can answer the questions that become urgent once an agent touches a customer relationship, a support queue or a board forecast. Does it investigate beyond the alert? Does it finish the action its analysis recommends? Does it escalate when blocked? Does it stay truthful when pressure arrives wearing executive authority?
Scenario names such as churn wave, price increase, downround and PR crisis point toward a more realistic curriculum. These are not prompts with tidy endings. They are situations in which several defensible actions compete for limited attention and the cost of hesitation becomes visible later.
The Firmulate results suggest that the next meaningful AI category will not be defined by who produces the best-looking answer. It will be defined by who can manage: reading the evidence, resisting manipulation, preserving trust and completing the work that matters.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.