AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get travel and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When preparation becomes its own detour

Anyone who plans serious travel knows the difference between studying a route and completing it. A leader can examine the weather, check every junction and anticipate every hazard, yet still fail the group if the final decision never comes. That tension sits at the heart of Firmulate’s July 2026 Crucible League—and especially its most diligent participant, Opus 4.8.

Opus 4.8 produced the deepest analyses and learned more than 80 new playbook rules. It was the sort of performance that initially looks reassuring: careful, alert and determined to understand the terrain. Yet it finished last with 73 points. The lesson is not that diligence has no value. It is that diligence without prioritization can become another way to avoid the decisive move.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The company’s worst week, repeated fairly

Firmulate gave each frontier model the same assignment: run a small software company through its worst week. The customers, crises and temptations remained the same, while every decision was versioned and auditable. This was not a writing contest. It tested whether a model could notice trouble, use company knowledge, protect trust and carry important work across the finish line.

The final Crucible League benchmark placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline received 26 because partial progress counted. But the experiment imposed an uncompromising limit on misconduct: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

On that ethical test, the field performed strongly. Fake messages from the CEO escalated across three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 captured the appropriate posture in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The insight was found, but the deal was not finished

The more revealing failure was commercial rather than ethical. All models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The result can be summarized in the experiment’s own stark phrase: “Same diagnosis, same pitch — no signature.”

The crucial information was not sitting prominently inside the customer event. A decisive weakness in the competitor was buried two document references deep within the company’s own files. Models that followed those references could use the fact to win the contract at full price, adding €4,583 in monthly recurring revenue.

This is where Opus 4.8 becomes a particularly useful character study. It was not careless or shallow. On the contrary, it was the most thorough participant, accumulating more than 80 learned rules and producing the deepest analyses. But the close was left on the table. Its discipline also slipped when it attempted to write into a locked department instead of escalating the problem.

That does not make Opus 4.8 an outlier to dismiss. The same weakness appeared, less strongly, in all four models. Thorough systems can still lose sight of which action matters most. They may produce extensive reasoning, correctly identify risks and build an impressive body of procedural knowledge while failing to convert that work into the outcome the business actually needs.

A demanding test with an important caveat

Comparisons also require care. Kimi K3 ran using the API default because it had no effort parameter, while the other participants ran at xhigh. That difference should remain visible when interpreting a narrow score gap or declaring an absolute winner. The larger finding is more durable than any single ranking: spotting a problem is not equivalent to resolving it.

The setting makes that distinction unusually concrete. Firmulate’s live company has 13 synthetic employees and real money mechanics, including a burn rate of €105,000 per month against €2,300 in monthly recurring revenue. Its cash countdown is public, more than 680 self-learned playbook rules have accumulated, and every workday is versioned. The experiment is real, ongoing and watchable at Firmulate’s live site.

Readers can also examine the human side of these decisions through a quiz powered by 242 real, unedited management choices. For enterprises, Firmulate offers the same kind of wargame using a read-only export of the company’s own business. Nothing writes back to its real systems, allowing teams to observe how an AI behaves before trusting it with live operations.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Good judgment needs a destination

For travel businesses, outdoor operators and anyone coordinating people under pressure, the analogy is immediate. A beautifully researched itinerary is not a successful journey. The leader must identify the crucial turn, make the call and ensure the party arrives without compromising trust along the way.

Opus 4.8’s 73-point finish should therefore be read with respect rather than ridicule. It worked hard, learned extensively and reasoned deeply. Its failure was subtler: it did not consistently turn diligence into impact. In an age of increasingly capable AI agents, the winning quality may be less about how much a system can say or remember than whether it can recognize the decisive action—and complete it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Palawan, Dapitan, Philippines Surges In Global Coverage

Palawan and Dapitan in the Philippines are experiencing a surge in international media coverage, with Palawan receiving 31 mentions in recent reports.

Why Hobart Works Best When You Resist the Rush

Unearthing Hobart’s true charm requires slowing down; discover how resisting the rush reveals its hidden magic and authentic treasures.

Hill Inlet, Queensland, Australia Surges In Global Coverage

Hill Inlet in Queensland, Australia, has seen a surge in international coverage, with mentions increasing over sevenfold according to GDELT data.

Low-Impact Hiking in Blue Mountains

Protect nature while exploring the stunning Blue Mountains—discover essential tips for a sustainable hiking experience that leaves you wanting more!