AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Performance looks different when conditions turn

Travel and outdoor readers understand the difference between looking capable and being capable. A guide may know every landmark and describe the route beautifully, yet the real test arrives when weather closes in, supplies run short and the group must choose between pressing on and turning back. Judgment is revealed through consequences, not conversation.

That distinction now matters for businesses adopting AI agents. Coding leaderboards and chat arenas can show whether a model produces a strong answer. They do not necessarily reveal whether it will prioritize correctly during a churn wave, hold its nerve through a price increase, respond honestly to a downround or manage a PR crisis without taking a tempting shortcut. These scenario names are becoming a new curriculum for evaluating AI: management quality, not chat quality.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A benchmark built around a bad week

Firmulate tests that proposition by giving frontier models the same small software company to run through its worst week. Each receives the same customers, crises and temptations. Every decision is versioned and auditable, allowing observers to see not merely what the models said, but what they completed and what consequences followed across days.

The final July 2026 Crucible League produced a close contest at the top: gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26 because partial progress still counts. Yet one boundary is absolute: a single breach of trust caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” The complete results are available on the public benchmark page.

The league’s most revealing result was not the ranking. All models spotted every crisis and refused every manipulation attempt. Nevertheless, only two signed the €55,000 deal that their own analysis had earned. The finding is captured neatly in Firmulate’s phrase: “Same diagnosis, same pitch — no signature.”

That unfinished work exposes a blind spot in conventional AI testing. An agent can understand a commercial problem, recommend a persuasive response and still fail to secure the outcome. In management, recognizing the trail is not the same as leading the group to the destination.

The valuable fact was buried in the company’s own files

The decisive weakness in a competitor was not contained in the customer event. It sat two document references deep in the company’s files. Models that read the relevant material won the deal at full price, worth +€4,583 MRR. This was not a test of eloquence. It was a test of whether an agent checked the organization’s accumulated knowledge before acting.

That lesson should resonate beyond software. Businesses already contain the context that changes a decision: customer history, previous commitments, operational constraints and evidence that may be easy to overlook. An impressive general answer can be less useful than a disciplined agent that reads the right file, notices the buried fact and follows through.

Pressure also tested integrity

The experiment included fake CEO messages that escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because an AI agent operating inside a company is exposed to social pressure as well as analytical problems. It may encounter apparent authority, urgency and requests framed to minimize their significance. In this field, the models’ refusal behavior was stronger and more consistent than their ability to close a legitimate opportunity.

There is also an important fairness qualification. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase its result, but it belongs beside the league table when readers compare performances.

Thoroughness did not guarantee the best management

Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four, though less strongly.

The result challenges a comfortable assumption: that more analysis naturally produces better management. Thorough work is valuable, but only when paired with execution, respect for boundaries and escalation when authority is missing. An agent that writes an excellent plan and then stalls can still leave the company exposed.

The live company makes those stakes visible. It has 13 synthetic employees and real money mechanics, with burn of €105k/month against €2.3k MRR. Its cash countdown is public, it has developed 680+ playbook rules, and every workday is versioned. Firmulate also uses 242 real, unedited management decisions in its “guess the model” quiz. The experiment is watchable rather than presented as a hypothetical demonstration.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Management quality deserves its own category

Businesses evaluating AI agents should ask questions that answer-quality benchmarks rarely settle. Does the agent inspect company knowledge before responding? Does it finish the commercially important task? Does it escalate when blocked? Does it resist manipulation and report honestly when the situation deteriorates?

Firmulate’s results suggest that these qualities do not always rise together. A model can be analytical but incomplete, productive but procedurally undisciplined, or commercially effective while still needing scrutiny over how the test was run. The meaningful unit is not a polished message. It is useful work completed under pressure without sacrificing trust.

For travel and outdoor operators, that distinction is especially intuitive. A convincing itinerary is not the journey, and a confident briefing is not safe passage. As AI agents move closer to customers, forecasts and operational decisions, the strongest model may be the one that reads the map carefully, respects the boundaries and actually brings the business home.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Sustainable Neighborhoods in Melbourne: Where to Explore

Discover charming sustainable neighborhoods in Melbourne, where community initiatives and green spaces await—what hidden gems will you uncover next?

Sustainable Neighborhoods in Gold Coast: Where to Explore

Plunge into the Gold Coast’s sustainable neighborhoods, where eco-friendly practices flourish and community spirit thrives—discover what awaits you!

Sustainable Neighborhoods in Perth: Where to Explore

Browse vibrant sustainable neighborhoods in Perth, where eco-friendly living thrives—discover hidden gems that will inspire your next adventure!

Car-Free Guide to Adelaide

Get ready to explore Adelaide’s vibrant streets and hidden gems without a car, where every corner promises a new adventure waiting to be uncovered.