firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

A business story that rewrites itself every workday

For news and media readers, Firmulate offers an unusually transparent corporate drama: a software company with 13 synthetic employees, real money mechanics and no shortage of jeopardy. It is burning €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown tracks the consequences.

This is not a retrospective case study polished for publication. The company runs every business day, versions each workday and makes its continuing fight for survival watchable online. Its synthetic staff have accumulated more than 680 self-learned playbook rules, and visitors can also read what the employees say. The result resembles a continuously updated business beat, except the company itself is producing the source material.

Amazon

AI management decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Extreme build-in-public

Build-in-public projects usually reveal selected milestones: a product launch, a revenue chart or a founder’s account of a difficult decision. Firmulate pushes the idea further by exposing an operating company’s daily struggle. Its income, burn and remaining time are not background details; they create the pressure under which every decision is made.

That pressure matters because Firmulate is designed to test whether frontier AI models can manage a company rather than merely discuss management. In the Crucible League, each model faced the same small software business during its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable.

The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. But the benchmark imposed a hard limit after any breach of trust, based on the principle that “no amount of good work outweighs a breach of trust.”

Seeing the crisis was not enough

Every model identified every crisis and rejected every manipulation attempt. The major difference emerged at the point where analysis had to become action: only two models signed the €55,000 deal that their own work had earned. The experiment’s blunt summary was: “Same diagnosis, same pitch — no signature.”

The detail that separated the successful performances was not sitting in the customer event. A decisive competitor weakness was buried two document references deep inside the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue.

That makes the experiment particularly relevant beyond the familiar question of whether an AI produces convincing prose. A model can identify a problem, prepare a credible response and still fail to complete the commercial task. In a company burning €105k each month against €2.3k MRR, that gap has immediate consequences.

Trust held under pressure

The models also encountered fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 refused. Kimi K3 recorded the clearest interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s result carries a fairness qualification. It ran with the API default and without an effort parameter, while the other models ran at xhigh. Even so, it finished behind gpt-5.6-sol by only the scores reported in the final league table and was one of the models that completed the deal.

Thoroughness did not guarantee success

Opus 4.8 illustrates another uncomfortable finding. It was the most thorough participant, producing the deepest analyses and adding 80 learned rules, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same discipline problem appeared in all four other participants, though less strongly.

That profile challenges a common assumption about capable AI: more analysis is not automatically better management. In Firmulate’s worst-week test, completion, attention to company files and procedural discipline mattered alongside diagnosis.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

A company and a continuing public record

Firmulate’s appeal as a media story is its refusal to freeze the experiment at a convenient conclusion. The league results provide a comparative snapshot, but the live company keeps operating under the harsher arithmetic of €105k monthly burn and €2.3k MRR.

Its public cash countdown gives the narrative a visible clock. Its 13 synthetic employees leave behind decisions and remarks. Its playbook of more than 680 learned rules records accumulated experience, while each workday becomes another version of the company.

The broader lesson is not that AI managers either succeed or fail in a single dramatic moment. It is that a company’s fate can turn on quieter acts: opening the referenced file, resisting a plausible impersonation, escalating when access is blocked and signing the deal after earning it. Firmulate has made those acts part of a real, ongoing and publicly observable business story.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Induction Cooking Secrets: The Pan Compatibility Rule Everyone Misses

Guess what crucial pan compatibility rule many overlook that can make or break your induction cooking success.

The €55,000 Test: Which AI Models Followed the Paper Trail?

Firmulate found that every AI saw the crisis, but only models that followed a buried document trail won the €55,000 deal at full price. It is a buying signal.

Smoker Temperatures: The “Stall” Explained Without the Mystery

Unlock the secrets of the stall and learn how maintaining steady smoker temperatures can transform your barbecue results.

Dutch Oven Weight Matters: Here’s Why Your Arms Will Notice

Pondering the impact of Dutch oven weight reveals why your arms will notice—uncover the key factors affecting handling and performance.