firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Can AI Save a Struggling Business? Watch It Live in Action

Imagine a company with no employees, losing €105,000 every month against a revenue of just €2,300. Now, add a twist: an AI-driven simulation where you can watch artificial managers handle crises, make decisions, and attempt to close deals — all in real time. This is not science fiction; it’s the live experiment from Firmulate, where AI models are tested as full-fledged companies, battling to survive against the odds.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Company as a Public Laboratory

At the heart of this experiment is a small software firm run entirely by AI models, which are tested against real-world challenges. The company has 13 synthetic employees, each governed by more than 680 self-learned rules, with every workday versioned and publicly viewable at firmulate.com/live.html. The firm faces continuous crises, customer negotiations, and the temptation to manipulate or cut corners. Its financial position is stark: burning €105,000 monthly while generating only €2,300 in recurring revenue, with a public cash countdown exposing its precarious situation.

The core idea is to see how different AI models perform under the same stressful conditions. Four frontier models, including GPT-5.6, Kimi K3, Sonnet 5, and Fable 5, were subjected to the same worst-week scenario, with their decisions recorded and auditable. Despite facing identical challenges, their outcomes varied significantly.

Key Findings: Crisis Detection and Integrity

All four models successfully identified every crisis — from customer complaints to internal threats. They refused every manipulation attempt, including social engineering tactics such as fake CEO messages and reporter trick questions. For example, all five models refused to sign a dubious €55,000 deal after analysis, citing risks and integrity concerns, demonstrating a high level of discipline.

Interestingly, the decisive advantage came from reading deeper into company files. A hidden document reference in the firm’s internal files contained crucial information, which only models that examined these references successfully closed the deal at full price (+€4,583 MRR). This shows that thorough information retrieval, not surface-level analysis, can be the decisive factor in AI decision-making.

Real Money Mechanics and the Build-in-Public Approach

The live experiment is transparent and continuously updated, with every decision logged and versioned daily. The company burns €105,000 each month against a mere €2,300 in revenue, highlighting the brutal reality of AI-driven business management under pressure. This open approach — the ‘build-in-public’ model — offers a rare window into how AI manages complex, real-world business dilemmas.

Among the models, OPUS 4.8 demonstrated the deepest analysis with over 80 learned rules and thorough reviews. Yet, it finished last, hesitating on closing a deal and slipping in discipline by leaving an approved deal unexecuted. This underscores that even deep analysis doesn’t guarantee success if discipline falters under stress.

What This Means for the Future of AI in Business

The experiment’s takeaway is clear: AI systems can identify crises and refuse manipulative tactics reliably, but their ability to execute deals or follow through with discipline varies. The real challenge lies in aligning AI behavior with business objectives and ensuring its decisions are both honest and effective.

For decision-makers, the question is no longer simply about AI’s ability to generate convincing responses but whether it can truly finish what it starts, read relevant information thoroughly, and uphold integrity when under pressure. As AI models advance, their role in managing customer relationships, support, and strategic decisions will depend heavily on these qualities.

Watch the Experiment Unfold

Visitors can watch this ongoing experiment in real time at firmulate.com/live.html. The site rebuilds twice a day, showcasing the latest decision-making performance of each model. This level of transparency is unprecedented in AI testing, offering a raw and unfiltered view of how artificial decision-makers operate in complex environments.

The company’s results are publicly available, with detailed scores: GPT-5.6 leads with 95 points, having uncovered a hidden fact and secured the deal; Kimi K3 follows closely with 93, executing the deal with disciplined integrity; Sonnet 5 scored 88, with minor slips; and Fable 5 scored 77, leaving potential on the table.

This experiment isn’t just about AI benchmarks; it’s a glimpse into the future of business management, where AI agents may one day run entire companies, or at least assist human managers in making more honest, consistent, and well-informed decisions.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Key Takeaway: AI Can Detect and Avoid Crises — But Discipline and Thoroughness Are Still Critical

This live experiment demonstrates that AI models can identify crises, refuse manipulative tactics, and even find hidden key information. However, success depends on their discipline and depth of analysis. The real-world implications are profound: AI’s role in business will hinge not only on what it knows but how reliably it executes and maintains integrity under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Pizza Ovens: The Temperature Numbers That Make or Break Crust

When it comes to pizza ovens, understanding the temperature ranges that define perfect crusts can make all the difference—discover how to master heat for your ideal slice.

Portable Induction Cooktops: The Wattage Math That Predicts Performance

Boost your cooking skills by understanding how wattage influences portable induction cooktop performance—discover the key factors that could change your kitchen game.

AI Showdown Reveals Hidden Weaknesses in Business Decision-Making

AI models can diagnose crises and resist manipulation, but only those that read internal data and follow through can close real deals. Live tests reveal a hidden gap in AI management skill.

Bread Makers Demystified: Settings That Actually Change Results

What you choose to adjust in your bread maker can dramatically change your loaf’s outcome, and understanding these settings is key to perfect results.