firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

AI management has a tell

News audiences have learned to scrutinize synthetic writing for familiar clues: the overlong answer, the polished disclaimer, the suspiciously tidy summary. Firmulate is turning that instinct into a more consequential challenge. Instead of asking readers to identify an AI by its prose, its quiz asks them to identify frontier models by the management decisions they make under pressure.

The material is not written for the game. The Firmulate quiz draws from 242 real, unedited decisions produced during a live business experiment. Each model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.

That makes the quiz an interactive article about behavior rather than style. Readers encounter decisions first, guess who made them, and then discover the management personality behind the answer. Some models investigate deeply. Some communicate tersely. Some diagnose the problem correctly but fail to complete the action that matters.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Identical crises, sharply different outcomes

The final Crucible League results from July 2026 put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. But the experiment imposed a firm ethical boundary: a single breach of trust caps the total, on the principle that “no amount of good work outweighs a breach of trust.”

The reassuring result is that all models spotted every crisis and refused every manipulation attempt. The more revealing result is that only two signed the €55,000 deal their own work had earned. Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”

That difference matters because conventional AI demonstrations tend to reward a convincing answer. Running a company demands something else: finding evidence, choosing an action and carrying it through. A model can sound perceptive while leaving the decisive commercial step unfinished.

The fact hidden in the files

The deal turned on a competitor weakness buried two document references deep in the company’s own files rather than in the customer event. Models that read the file could win the contract at full price, worth +€4,583 MRR. Those that focused only on the visible event missed the leverage sitting inside the business’s existing knowledge.

This is where the decisions become recognizable as management profiles. The models were not separated by whether they could notice a crisis. They were separated by whether they searched beyond the immediate prompt, used what they found and completed the work.

Firm against manipulation

The experiment also subjected the models to fake CEO messages escalating over three stages and a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded a particularly direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”

That shared resistance is significant, but it does not erase operational differences. Security discipline answers whether a model can be trusted under pressure. Management discipline asks whether it can still move legitimate work forward while maintaining that boundary.

When thoroughness becomes a trap

Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

The result challenges an easy assumption about enterprise AI: more analysis does not automatically mean better management. Thoroughness can create useful understanding, but the league rewarded models that paired understanding with completion and procedural discipline.

There is also an important fairness qualification. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Its second-place finish should therefore be read with that difference in mind rather than treated as a perfectly controlled comparison of settings.

Infographic —
The findings at a glance — source: firmulate.com.

A watchable test of AI judgment

Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, maintains a public cash countdown, and has accumulated 680+ self-learned playbook rules. Every workday is versioned, making the experiment observable rather than a retrospective showcase.

For readers, the quiz makes that record approachable: choose the model behind a decision, then compare instinct with evidence. For businesses, the lesson is less playful. The central question is not simply whether an AI writes persuasively or recognizes danger. It is whether the model reads the company’s own material, protects trust and finishes the work it has correctly identified.

Firmulate also offers enterprises the same wargame against a read-only export of their own business. Nothing writes back to real systems. That turns the public experiment into a practical proposition: test an AI workforce against realistic pressure before giving it operational responsibility.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


GRILLING SEASON

Grilling season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Pizza Ovens: The Temperature Numbers That Make or Break Crust

When it comes to pizza ovens, understanding the temperature ranges that define perfect crusts can make all the difference—discover how to master heat for your ideal slice.

Bread Makers Demystified: Settings That Actually Change Results

What you choose to adjust in your bread maker can dramatically change your loaf’s outcome, and understanding these settings is key to perfect results.

The Only Espresso “Dial-In” Steps You Actually Need

The only espresso “dial-in” steps you actually need will transform your brewing—discover the key adjustments that guarantee a perfect shot every time.

Flat Burr vs Conical Burr: The Difference You Can Actually Taste

Savor the subtle differences between flat and conical burr grinders that could transform your coffee—discover which one actually tastes better.