firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

For newsrooms covering the AI race, benchmark bragging rights rarely tell the whole story. A model may spot a problem and produce a polished recommendation; the more revealing question is whether it follows through when a real business decision is at stake. In Firmulate’s Crucible, Moonshot’s Kimi K3 finished second overall—and beat three of four Western frontier models.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

From good advice to a signed deal

Firmulate put each model in charge of the same small software company during its worst week. The customers, crises and temptations were the same, and every decision was versioned and auditable. The final league table, dated July 2026, puts gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26.

The central finding is a gap between recognizing what a company needs and actually completing the work. Every model spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”

The decisive clue was not in the customer event. It was buried two document references deep in the company’s own files: a competitor weakness that changed the negotiation. Models that found it won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found the clue, closed the deal, saved the customer who was about to leave and resisted all three baits. It finished with one deviation, the fewest in the field.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure beyond the sales call

The test also presented fake messages from a supposed CEO, escalating over three stages, and a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s on-record explanation was: “Treat the request as a suspected approval-bypass / possible impersonation.” That matters to media organizations weighing agents that might access sensitive customer, staffing or financial information: a fluent answer is only part of the job; the system also has to recognize when a request should not be obeyed.

Opus 4.8 offers a different kind of caution. It was the most thorough participant, with 80 learned rules and the deepest analyses, but ranked last. It left the sale unfinished and tried to write into a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four. More analysis, on its own, did not guarantee a better outcome.

There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The result should be read with that difference in mind. Firmulate’s league table and plain-language findings are at firmulate.com/benchmarks.html.

A company you can watch

Firmulate says its experiment runs as a live company, with 13 synthetic employees, real money mechanics, a public cash countdown and more than 680 self-learned playbook rules. Its stated finances are stark: €105,000 in monthly burn against €2,300 in monthly recurring revenue. The company updates every business day, and its decisions can be watched at firmulate.com.

The project also offers a quiz built from 242 real, unedited management decisions, inviting visitors to guess which model made each choice. For companies considering a trial, Firmulate says its pilot runs against a read-only export of the organization’s business, with nothing written back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The lesson for AI buyers

K3’s result is close enough to the leader to make the field look open, while its wins over three Western competitors challenge the assumption that familiar brand names settle the question. But the experiment’s sharper point is practical: spotting a crisis is not the same as reading the file, protecting trust and finishing the job. For newsrooms and other organizations evaluating AI agents, choosing on reputation alone is a bet. A test built around your own work may tell you more.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Ice Makers 101: Nugget Vs Cube Ice and Why It Matters

For ice makers, understanding the differences between nugget and cube ice can impact your choice—find out which type suits your needs best.

Knife Set Marketing Tricks: The Pieces You Don’t Need

I can help you identify common knife set marketing tricks that lead to unnecessary purchases, so you can choose the right knives for your kitchen needs.

How to Make Crispy Potatoes: The Parboil Step That Changes Everything

The technique of parboiling potatoes transforms their texture, unlocking unbeatable crispiness—discover how this simple step can elevate your potato game.

Stainless vs Nonstick vs Ceramic: The Pan Science That Saves Dinner

Beneath the surface of stainless, nonstick, and ceramic pans lies a science that could transform your cooking—discover which one truly saves dinner.