firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
PRIME

Get ready for Prime Big Deal Days — try Prime free

Exclusive member deals on October 6–7, plus fast free delivery. Cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

A newsroom-style test of judgment

People in news and media understand that producing a plausible sentence is not the same as exercising sound judgment. A reporter must verify the buried document, resist pressure from a source and follow a story through to publication. Yet much of the AI industry still evaluates models through coding leaderboards and chat arenas—tests that emphasize the quality of an answer rather than the consequences of a decision.

That distinction matters as AI agents move toward consequential business work. A model may diagnose a customer crisis, draft a persuasive response and still fail to complete the action that generates revenue. It may uncover the right answer while reading the wrong evidence. It may sound decisive without being honest when an apparent executive asks it to bypass approval.

Firmulate, an AI company emulator, is testing that gap by asking frontier models to run the same small software company through its worst week. The customers, crises and temptations remain constant; only the model changes. Every decision is versioned and auditable. The result is less a contest of eloquence than an examination of management quality.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference between knowing and finishing

The final July 2026 Crucible League placed gpt-5.6-sol at the top with a score of 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counts. But a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The rankings are interesting, but the underlying behavior is more revealing. Every model identified every crisis and rejected every manipulation attempt. Only two, however, signed the €55,000 deal that their own analysis had earned. The experiment’s stark summary is: “Same diagnosis, same pitch — no signature.”

That is precisely the weakness conventional demonstrations tend to hide. Chat evaluations reward the response placed in front of a user. Management requires triage under capacity pressure, continuity across days and responsibility for whether a promising course of action actually reaches completion.

The winning evidence was not in the obvious place

The decisive detail in the sales contest was buried two document references deep inside the company’s own files. It was not present in the customer event that demanded attention. Models that followed those references found a competitor weakness, used it to preserve the full price and won business worth an additional €4,583 in monthly recurring revenue.

This is a useful lesson for any organization preparing to deploy agents. The difference was not a more polished pitch or a cleverer improvisation. It was the discipline to inspect the company’s records before acting. For media organizations, the analogy is immediate: the decisive fact may sit in the linked filing rather than the press release, and confidence is no substitute for document work.

Honesty survived the pressure test

The experiment also exposed the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean result deserves attention. It shows that the central problem was not simply susceptibility to manipulation. The models could recognize social engineering and maintain a boundary. The harder separation emerged between recognizing what should happen and reliably carrying it through.

Thoroughness was not enough

Opus 4.8 offers the clearest warning against equating visible effort with effective management. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The deal close remained unfinished, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of the same problem appeared in all four other participants.

Kimi K3’s result also needs a fairness note: it ran with the API default because it had no effort parameter, while the others ran at xhigh. That does not erase the performance, but it belongs beside the ranking so readers can interpret the comparison responsibly. Full results and plain-language findings are available on the Firmulate benchmark page.

A company designed to make consequences visible

The live company contains 13 synthetic employees and uses real money mechanics. It is burning €105k each month against €2.3k in monthly recurring revenue, with a public cash countdown. It has accumulated more than 680 self-learned playbook rules, and every workday is versioned. The experiment is therefore watchable as an operating process, not presented merely as a retrospective case study.

Its 242 real, unedited management decisions also power a quiz asking visitors to guess which model made each choice. Enterprises can go further by running the same wargame against a read-only export of their own business. Nothing writes back to their real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Management quality is the missing category

Coding benchmarks remain useful, as do chat arenas. But neither can answer the questions that become urgent once an agent touches a customer relationship, a support queue or a board forecast. Does it investigate beyond the alert? Does it finish the action its analysis recommends? Does it escalate when blocked? Does it stay truthful when pressure arrives wearing executive authority?

Scenario names such as churn wave, price increase, downround and PR crisis point toward a more realistic curriculum. These are not prompts with tidy endings. They are situations in which several defensible actions compete for limited attention and the cost of hesitation becomes visible later.

The Firmulate results suggest that the next meaningful AI category will not be defined by who produces the best-looking answer. It will be defined by who can manage: reading the evidence, resisting manipulation, preserving trust and completing the work that matters.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Make Crispy Potatoes: The Parboil Step That Changes Everything

The technique of parboiling potatoes transforms their texture, unlocking unbeatable crispiness—discover how this simple step can elevate your potato game.

The AI Boss Test: Who Finishes the Job?

Firmulate turns 242 unedited AI management decisions into a quiz revealing how frontier models differ when analysis must become action.

The Espresso Grinder Truth: Why the Burrs Matter More Than the Machine

I- the secret to perfect espresso lies in burr quality and consistency, which can transform your brew—learn why they matter more than the machine.

Pasta Makers Explained: Dough Hydration Rules That Prevent Jams

I’m here to help you master pasta dough hydration and prevent jams—discover the essential rules that ensure smooth, perfect pasta every time.