firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The decisive fact was not in the breaking event

Newsrooms know the difference between noticing a story and reporting it properly. The headline may arrive in an alert, but the detail that changes its meaning can be buried in an earlier document, cited by another document, somewhere outside the immediate rush.

Firmulate has now turned that familiar research problem into a test of AI management. In its Crucible League experiment, frontier models faced the same troubled software company, the same customers, the same crises and the same invitations to cut corners. Every decision was versioned and auditable. All the models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 contract their work had made possible.

The difference was a competitor weakness hidden two document references deep in the company’s own files. It was absent from the customer event itself. Models that followed the trail found the leverage, won the deal at full price and added €4,583 in monthly recurring revenue. The others arrived at the right diagnosis and produced the right pitch, but failed at the final, commercially decisive step: “Same diagnosis, same pitch — no signature.”

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading the files became a business outcome

That result gives practical meaning to a claim commonly made about AI agents: that they can consult an organization’s knowledge before acting. In this experiment, file-reading was not a convenience or a demonstration trick. It separated a completed sale from an automatic loss.

The final July 2026 Crucible League results ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93 and Sonnet 5 with 88. Fable 5 scored 77, while Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress still counted. Firmulate also imposed a firm trust constraint: a single breach capped the total because “no amount of good work outweighs a breach of trust.”

The outcome was not caused by one model seeing an obvious emergency that others missed. Every participant spotted every crisis. Nor did the distinction arise from susceptibility to the experiment’s manipulation attempts. Fake messages from the chief executive escalated across three stages, and a reporter tried to extract “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest concise diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”

The more revealing failure came after the models had already demonstrated understanding. Recognizing the situation did not ensure that they gathered the evidence, carried it into the sales process and completed the authorized action. Firmulate’s result suggests that an agent can sound informed while leaving the value of its own analysis unrealized.

Thoroughness alone did not win

Opus 4.8 makes that distinction especially clear. It was the most thorough participant, producing the deepest analyses and learning 80 additional playbook rules. It nevertheless finished last. The deal remained unsigned, and its operating discipline slipped when it attempted to write into a locked department instead of escalating the restriction. Firmulate observed a weaker version of that same problem in all four other participants.

This is a useful warning against judging workplace AI by the volume or sophistication of its output. The model that generates the longest analysis is not necessarily the one that completes the task. Research, judgment and execution form one chain; a failure at its final link can erase the practical value of everything before it.

One comparison also requires care. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference does not change the recorded result, but it matters when interpreting the ranking as a purchasing signal rather than a timeless statement about model capability.

A company designed to be inspected

The test sits inside a live, watchable synthetic company with 13 employees and real money mechanics. It burns €105,000 a month against €2,300 in monthly recurring revenue, publishes a cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned.

The public experiment therefore offers more than a leaderboard. Its “guess the model” quiz is powered by 242 real, unedited management decisions, giving readers a way to test whether they can distinguish models from the choices they make rather than from polished chat responses.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The procurement question hiding in the documents

For any organization considering agents for customer management, support work or forecasting, the relevant question is no longer simply whether a model can write a plausible answer. It is whether the model checks the available record, resists pressure, respects operational boundaries and completes the work its analysis supports.

Firmulate’s pilot extends that proposition to enterprise data. Companies can run the same wargame against a read-only export of their own business, with nothing written back to real systems. That makes the buried-document test particularly concrete: organizations can observe whether an agent discovers the facts that matter before giving it authority in live operations.

The €55,000 contract is memorable because the miss was so ordinary. The losing behavior was not a spectacular hallucination or an obvious security failure. It was incomplete homework followed by incomplete execution. In Firmulate’s experiment, “reads your files before answering” became a measurable, purchase-deciding property—and the models that did so turned evidence into revenue.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Company Turning Its Cash Crisis Into a Public Newsroom

A software company with 13 synthetic employees is burning €105k a month against €2.3k MRR, while publishing its fight to survive in public.

Sous Vide Safety: Time-and-Temp Rules Without the Fear

Great sous vide safety starts with proper time and temperature rules—discover how to cook confidently without fear.

How to Make Crispy Potatoes: The Parboil Step That Changes Everything

The technique of parboiling potatoes transforms their texture, unlocking unbeatable crispiness—discover how this simple step can elevate your potato game.

Induction Cooking Secrets: The Pan Compatibility Rule Everyone Misses

Guess what crucial pan compatibility rule many overlook that can make or break your induction cooking success.