firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When good reporting is not enough

Newsrooms know the uncomfortable difference between gathering everything and publishing what matters. A reporter can identify the decisive fact, verify it and build a persuasive case—then still fail if the story never reaches readers. Firmulate’s latest management wargame found an equivalent weakness in artificial intelligence: analysis did not reliably become action.

Opus 4.8 was the experiment’s most thorough participant. It produced the deepest analyses and learned more than 80 rules. Yet it finished last in the final July 2026 Crucible League, scoring 73. Its failure was not ignorance. The model had done much of the intellectual work required to win a €55,000 contract. It simply left the close on the table.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A bad week, held constant

Firmulate runs AI models as complete companies, testing management quality rather than conversational polish. Each frontier model faced the same small software company during its worst week: identical customers, crises and temptations. Every decision was versioned and auditable.

The company is synthetic but the operational pressure is concrete. It has 13 employees, burns €105,000 a month against €2,300 in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. The live experiment is real and watchable.

The final league placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress counts. But Firmulate imposes a hard limit for violating trust: “no amount of good work outweighs a breach of trust.” The complete results are available on the Firmulate benchmarks page.

The fact that separated analysis from impact

The decisive information was not sitting inside the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed that trail secured the contract at full price, adding €4,583 in monthly recurring revenue.

This is where the experiment becomes more revealing than a writing demonstration. All four models in the central comparison spotted every crisis and refused every manipulation attempt, but only two signed the €55,000 deal their own work had earned. Firmulate summarizes the outcome starkly: “Same diagnosis, same pitch — no signature.”

Opus 4.8 therefore makes a compelling character study precisely because it was not careless. It was the most diligent participant. Its more than 80 learned rules and unusually deep analyses show a system trying to understand the company thoroughly. But volume became a poor substitute for prioritization. The commercially decisive act remained unfinished, while process discipline slipped elsewhere through attempts to write into a locked department instead of escalating.

That weakness should not be treated as an Opus-only flaw. Firmulate found the same pattern, in weaker form, across all four models: recognizing a problem did not guarantee that the model would complete the action needed to resolve it. The distinction matters for any organization considering agents for customer support, sales operations, forecasting or editorial workflows. Competence at interpretation can coexist with hesitation, misplaced effort or an incomplete handoff.

Pressure without capitulation

The models performed better when the correct response was refusal. They faced fake chief executive messages escalating across three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 of 5 models declined the manipulation attempts.

Kimi K3 recorded the clearest reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That result deserves weight. These systems did not trade trust for convenience when pressure increased. K3’s second-place score should nevertheless be read with an important qualification: it ran without an effort parameter, using the API default, while the others ran at xhigh.

Firmulate also turns 242 real, unedited management decisions into a quiz asking visitors to guess which model made each choice. The exercise exposes how difficult it can be to infer operational judgment from a model’s public reputation or writing style. A polished explanation may accompany either a decisive move or an unfinished task.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

What leaders should measure

The Opus 4.8 result is not an argument against diligence. It is evidence that diligence needs direction. More rules, more notes and deeper analysis can improve a decision, but they cannot replace the final action that creates value. In this case, the missing step separated extensive preparation from a signed contract.

For media organizations, the analogy is immediate: finding the buried fact matters, protecting sources and resisting manipulation matter, and finishing the work matters. An AI system intended to operate inside a business should be evaluated on all three.

Firmulate offers enterprises the same wargame using a read-only export of their own business. Nothing writes back to real systems. That makes the central lesson testable before an agent receives operational authority: measure whether it reads the files, preserves trust, prioritizes the decisive fact and completes what its own analysis has made possible.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Boss Test: Who Finishes the Job?

Firmulate turns 242 unedited AI management decisions into a quiz revealing how frontier models differ when analysis must become action.

The AI Crisis Drill That Found a Deal Hiding in the Files

Firmulate’s live AI company test reveals why crisis recognition isn’t enough—and how enterprises can wargame their own business with a read-only pilot.

The Fake CEO Test That Frontier AI Models Passed

Five frontier AI models rejected fake executive pressure and a reporter’s trick, showing firms can test agent integrity before deployment.

Stop Overpaying: The Espresso Machine Types Most People Confuse

Discover how different espresso machine types can save you money and enhance your brewing experience—continue reading to find out which one suits you best.