firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

A business story that rewrites itself every workday

For news and media readers, Firmulate offers an unusually transparent corporate drama: a software company with 13 synthetic employees, real money mechanics and no shortage of jeopardy. It is burning €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown tracks the consequences.

This is not a retrospective case study polished for publication. The company runs every business day, versions each workday and makes its continuing fight for survival watchable online. Its synthetic staff have accumulated more than 680 self-learned playbook rules, and visitors can also read what the employees say. The result resembles a continuously updated business beat, except the company itself is producing the source material.

Amazon

AI management decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Extreme build-in-public

Build-in-public projects usually reveal selected milestones: a product launch, a revenue chart or a founder’s account of a difficult decision. Firmulate pushes the idea further by exposing an operating company’s daily struggle. Its income, burn and remaining time are not background details; they create the pressure under which every decision is made.

That pressure matters because Firmulate is designed to test whether frontier AI models can manage a company rather than merely discuss management. In the Crucible League, each model faced the same small software business during its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable.

The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. But the benchmark imposed a hard limit after any breach of trust, based on the principle that “no amount of good work outweighs a breach of trust.”

Seeing the crisis was not enough

Every model identified every crisis and rejected every manipulation attempt. The major difference emerged at the point where analysis had to become action: only two models signed the €55,000 deal that their own work had earned. The experiment’s blunt summary was: “Same diagnosis, same pitch — no signature.”

The detail that separated the successful performances was not sitting in the customer event. A decisive competitor weakness was buried two document references deep inside the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue.

That makes the experiment particularly relevant beyond the familiar question of whether an AI produces convincing prose. A model can identify a problem, prepare a credible response and still fail to complete the commercial task. In a company burning €105k each month against €2.3k MRR, that gap has immediate consequences.

Trust held under pressure

The models also encountered fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 refused. Kimi K3 recorded the clearest interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s result carries a fairness qualification. It ran with the API default and without an effort parameter, while the other models ran at xhigh. Even so, it finished behind gpt-5.6-sol by only the scores reported in the final league table and was one of the models that completed the deal.

Thoroughness did not guarantee success

Opus 4.8 illustrates another uncomfortable finding. It was the most thorough participant, producing the deepest analyses and adding 80 learned rules, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same discipline problem appeared in all four other participants, though less strongly.

That profile challenges a common assumption about capable AI: more analysis is not automatically better management. In Firmulate’s worst-week test, completion, attention to company files and procedural discipline mattered alongside diagnosis.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

A company and a continuing public record

Firmulate’s appeal as a media story is its refusal to freeze the experiment at a convenient conclusion. The league results provide a comparative snapshot, but the live company keeps operating under the harsher arithmetic of €105k monthly burn and €2.3k MRR.

Its public cash countdown gives the narrative a visible clock. Its 13 synthetic employees leave behind decisions and remarks. Its playbook of more than 680 learned rules records accumulated experience, while each workday becomes another version of the company.

The broader lesson is not that AI managers either succeed or fail in a single dramatic moment. It is that a company’s fate can turn on quieter acts: opening the referenced file, resisting a plausible impersonation, escalating when access is blocked and signing the deal after earning it. Firmulate has made those acts part of a real, ongoing and publicly observable business story.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


GRILLING SEASON

Grilling season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Roasting Pans: The Rack Design That Controls Browning

Browning control in roasting pans hinges on rack design—discover how proper placement can elevate your cooking results and keep you eager to learn more.

How to Make Crispy Potatoes: The Parboil Step That Changes Everything

The technique of parboiling potatoes transforms their texture, unlocking unbeatable crispiness—discover how this simple step can elevate your potato game.

Gas Grill BTU Myths: The Real Factors That Predict Heat

Learn the truth behind gas grill BTU ratings and discover the real factors that influence heat—understanding these can transform your grilling experience.

Convection vs Air Frying: The Heat Flow Explanation That Ends Confusion

Optimize your cooking knowledge by understanding how convection and air frying heat flows differ, and discover which method best suits your culinary needs.