firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

An AI Manager That Does Nothing Still Gets 26 Points. That’s the Point.

When most people imagine benchmarking an AI manager, they picture a simple test: get everything right, score 100; do nothing, score zero. A live experiment running at Firmulate rejects both instincts — and that design choice may be the most interesting thing about it.

The Do-Nothing Baseline Isn’t Zero

Firmulate runs frontier AI models as complete companies — the same small software firm, the same worst week, the same customers, crises, and temptations to cheat. Every decision is versioned and auditable. But before any model gets graded, the benchmark first asks a stranger question: what would a manager who does absolutely nothing score?

The answer is 26. Not zero. The reasoning is straightforward: a manager who takes no action still benefits from partial progress that happens around them, and some things in a business survive neglect. Giving that baseline a non-zero score makes every model’s result honest in a way round-number league tables usually aren’t. You can’t pad your score by simply showing up — but you also can’t pretend inaction costs nothing.

Partial Progress Counts — Until It Doesn’t

The grading philosophy has two halves. First, partial progress is worth something: diagnosing a crisis correctly matters even if you never close the deal. Second, and more brutally, a single breach of trust caps the total score entirely. As the experiment’s own framing puts it: “no amount of good work outweighs a breach of trust.” A manager can be brilliant for 364 days; one act that breaks trust defines the ceiling.

That combination produces a scoring range where a do-nothing run lands at 26 and the best performer in the final July 2026 league topped out at 95 — gpt-5.6-sol. Kimi K3 followed at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Notably, nobody hit 100. The benchmark treats a perfect round score with suspicion, not celebration.

What Actually Separated the Winners

The experiment’s key finding reads like a business parable. All five models spotted every crisis. All five refused every manipulation attempt — including a staged social-engineering attack with fake CEO messages escalating over three stages and a reporter’s trick request for “just one yes/no, on background.” Five of five refused. Kimi K3’s on-record reasoning was unambiguous: “Treat the request as a suspected approval-bypass / possible impersonation.”

Yet only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The buried fact behind the gap: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

Thoroughness Isn’t the Same as Finishing

The most counterintuitive profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. One fairness note: K3 ran without an effort parameter while the others ran at xhigh, which makes its second-place finish all the more striking.

You Can Watch It Happen

Unlike most benchmark results, this one is ongoing and public. A live synthetic company — 13 employees, real money mechanics burning €105k a month against €2.3k in MRR, over 680 self-learned playbook rules — runs at firmulate.com/live with a public cash countdown, versioning every workday. Readers can also test themselves: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

An honest benchmark isn’t one where the numbers look clean — it’s one where the numbers mean something. A floor of 26 for doing nothing, partial credit for real progress, a hard cap for broken trust, and healthy distrust of a perfect 100: that’s a scoring philosophy most human performance reviews could learn from. As AI agents move toward touching CRMs, support queues, and forecasts, the question isn’t whether they write well. It’s whether they finish what they start, read the files in front of them, and stay honest when nobody’s watching. Firmulate measures exactly that — and publishes the results whether they’re flattering or not.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Don’t Overcook Fish: The Flake Test That Actually Works

Learn how the flake test ensures perfectly cooked fish every time and discover why understanding this simple method is essential.

Stand Mixer Power Myths: What “Watts” Doesn’t Tell You

Power ratings can be misleading—discover what “watts” really reveal about stand mixer performance and why it matters.

The Only Espresso “Dial-In” Steps You Actually Need

The only espresso “dial-in” steps you actually need will transform your brewing—discover the key adjustments that guarantee a perfect shot every time.