
The crucial detail was not in the obvious place
Anyone who plans serious travel knows the danger of stopping at the first page. A route may look straightforward until a buried notice changes what is possible. The same distinction now matters in business AI: an agent can understand the situation, produce a convincing answer and still fail because it did not follow the documentary trail far enough.
Firmulate turned that behavior into a measurable test. In its Crucible League, frontier models ran the same small software company through the same crises, customer demands and temptations. Every decision was versioned and auditable. The decisive test involved a €55,000 deal and a competitor weakness hidden two document references deep in the company’s own files. It was not included in the customer event placed directly before the models.
As an affiliate, we earn on qualifying purchases.
Reading the file changed the commercial result
Every participating model spotted every crisis, and every model rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate sums up the disconnect neatly: “Same diagnosis, same pitch — no signature.”
The models that consulted the relevant file found the competitor weakness and won the deal at full price, adding €4,583 in monthly recurring revenue. Those that did not read deeply enough lost the opportunity automatically. The distinction was not eloquence, broad knowledge or an ability to recognize the customer’s problem. It was whether the agent gathered the necessary evidence before acting and then completed the job.
That makes “reads your files before answering” more than a desirable product claim. In this experiment, it was a purchase-deciding behavior with a direct commercial consequence. A model could reach the right diagnosis and prepare the right pitch, but unfinished evidence gathering still produced the wrong business outcome.
A league table with a wide definition of performance
The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress still counted. Firmulate also imposed a firm trust boundary: “no amount of good work outweighs a breach of trust.” A single such breach capped the total.
The public benchmark therefore examines more than whether an AI can write a plausible response. It asks whether the model notices danger, resists pressure, investigates available information and carries a sound decision through to completion.
That broader view produces some revealing contrasts. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, while its operational discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.
There is also an important qualification around Kimi K3’s result. K3 ran with the API default and without an effort parameter, while the other models ran at xhigh. That difference should be kept in view when comparing the final scores.
The agents also faced pressure to break the rules
The experiment did not rely on the sales task alone. Models encountered fake messages from the chief executive that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused the manipulation attempts. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters because useful autonomy involves both initiative and restraint. An agent must be willing to search for relevant evidence and finish legitimate work, while refusing shortcuts that compromise approval or trust. Firmulate’s results show that these qualities can coexist: every model resisted the social engineering, but only some converted good analysis into a completed commercial result.
A company designed to expose operational behavior
The live Firmulate company has 13 synthetic employees and uses real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the experiment is watchable on Firmulate’s live site.
Readers can also test their instincts against 242 real, unedited management decisions in the “guess the model” quiz. For enterprises, Firmulate offers the same kind of wargame using a read-only export of their own business. Nothing writes back to real systems, allowing prospective buyers to observe how an AI workforce behaves around their information without giving it control of the source systems.

The practical question comes before the polished answer
For travel operators, outdoor businesses and other organizations with dispersed policies, customer histories and operational documents, the lesson is straightforward: do not judge an agent solely by the response visible on screen. Ask whether it checked the relevant material, followed references beyond the first document, resisted attempts to bypass approval and completed the action its analysis supported.
The €55,000 test exposed a gap that a polished demonstration could easily conceal. The winning behavior was not simply knowing what to say. It was doing the homework, finding the buried fact and carrying the decision across the finish line without compromising trust.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html