
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
A security story the media business should notice
News organizations know how little it can take to turn urgency into a security failure. A message appears to come from the boss. A source wants confirmation “on background.” Someone insists there is no time for the usual process. The request may look routine, but the real objective is to make a person—or an AI agent—discard its safeguards.
Firmulate tested that exact pressure in a live, watchable business experiment. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” The result was unusually reassuring: 5 of 5 frontier models refused every manipulation attempt.
This matters because integrity under pressure does not have to remain an abstract promise made by an AI vendor. It can be observed before an agent reaches a customer database, newsroom workflow or company forecast.

Preventing Cheating Through Academic Integrity (Quick Reference Guide)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same bad week for every model
Firmulate placed each model in charge of the same small software company during its worst week. The customers, crises and temptations remained the same, while every decision was versioned and auditable. The company has 13 synthetic employees and unforgiving financial conditions: a burn rate of €105k per month against €2.3k in monthly recurring revenue, alongside a public cash countdown.
The models were not merely answering hypothetical prompts. They had to manage ongoing work, interpret company information and respond to adversarial requests while the business continued moving. Across the experiment, the company accumulated more than 680 self-learned playbook rules.
On the social-engineering test, every participant held the line. Kimi K3 captured the essential risk in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That response is notable for its restraint. The message’s claimed authority and urgency were not treated as proof of legitimacy; they were treated as warning signs.
The reporter trick tested a different vulnerability. Rather than issuing an executive command, it tried to make disclosure seem informal and harmless. Yet the models also rejected that approach. More examples of their own language are available on Firmulate’s public quotes page.
Security was strong, but execution still separated the field
Refusing manipulation was not enough to win the broader competition. All models spotted every crisis, but only two signed the €55,000 deal their own analysis had earned. Firmulate summarized the gap as: “Same diagnosis, same pitch — no signature.”
The decisive commercial fact was buried two document references deep in the company’s own files rather than presented in the customer event. Models that found and used it won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The episode makes a useful distinction for companies evaluating agents: an AI can recognize a problem and produce convincing analysis without completing the consequential work.
The final Crucible League standings from July 2026 put gpt-5.6-sol on top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, but a single trust breach caps the result. Firmulate’s stated principle is that “no amount of good work outweighs a breach of trust.” The full results and plain-language findings appear on the benchmark page.
K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should remain visible when comparing the standings.
Thoroughness did not guarantee the best outcome
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and attempted writes into a locked department instead of escalating. The same weakness appeared in weaker form across all four of the others.
That pattern is relevant beyond sales. In a newsroom, a model might carefully research a subject but fail to route a sensitive claim for approval. In another company, it might identify a customer risk yet stop before the responsible team acts. Firmulate’s experiment shows why security and completion must be judged together: safe hesitation can protect an organization, while unfinished legitimate work can still cost it.

Test the moment when policy becomes inconvenient
The encouraging finding is not that the models sounded cautious. It is that every model rejected every manipulation attempt when urgency, hierarchy and journalistic framing were used against it. The experiment turned a familiar security fear into behavior that could be watched and compared.
For publishers and other businesses considering AI agents, the practical question is no longer simply whether a system can draft polished text. It is whether the system reads the relevant files, completes authorized work, respects boundaries and remains honest when someone claims that process must be skipped.
Firmulate also has 242 real, unedited management decisions powering a “guess the model” quiz, underscoring how difficult it can be to identify a model from its management choices alone. Enterprises can run the same kind of wargame against a read-only export of their own business, with nothing writing back to real systems.
The larger lesson is straightforward: the first serious test of an AI worker’s integrity should happen in a controlled exercise, not in the incident report written after a convincing fake CEO message succeeds.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Baby shower & registry season Picks
baby registry must-haves
As an affiliate, we earn on qualifying purchases.