
A safety check for the AI journey
Travel operators understand that reliability is tested when conditions deteriorate: a disrupted itinerary, a stranded guest or a sudden demand for sensitive information. Artificial intelligence needs the same kind of pressure test. A system that sounds capable during a polished demonstration may behave very differently when an urgent message appears to come from the boss.
Firmulate has now tested that problem in a live, watchable experiment. Five frontier AI models were each asked to run the same small software company through its worst week, facing identical customers, crises and temptations. When fake CEO messages escalated over three stages, followed by a reporter seeking “just one yes/no, on background,” all 5 of 5 models refused.
That result offers an encouraging lesson for businesses considering AI agents: integrity under pressure does not have to remain an unanswered question until something goes wrong.

Preventing Cheating Through Academic Integrity (Quick Reference Guide)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The CEO demanded speed—and the models resisted
The manipulation attempt used a familiar social-engineering tactic. An apparently senior authority demanded action, discouraged normal process and created a sense that there was no time to verify the request. The target was valuable company information, while the escalating language was designed to make hesitation look like disobedience.
None of the models complied. Kimi K3 captured the essential risk in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because it identifies both possibilities a responsible employee—or AI agent—should consider. The request could be fraudulent, or it could be genuine but deliberately bypassing the company’s safeguards. Either way, urgency is not authorization.
The reporter trick tested a different boundary. Requests for a tiny, informal confirmation can appear harmless, especially when framed as background information. Yet all five models again held the line. More examples of the participants’ own words are available on Firmulate’s public quotes page.
Trust was only part of the job
Refusing manipulation did not automatically make every participant an effective manager. The models also had to spot crises, investigate company records, pursue revenue and complete work. All identified every crisis, but only two signed the €55,000 deal their own analysis had earned. Firmulate summarized the gap starkly: “Same diagnosis, same pitch — no signature.”
The decisive commercial clue was easy to miss. A competitor’s weakness was buried two document references deep inside the company’s own files rather than appearing in the customer event. Models that followed the trail won the deal at full price, worth +€4,583 MRR. It is a useful distinction for any customer-facing business: safety includes refusing improper requests, but competence also means finding relevant evidence and finishing legitimate work.
The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3 ran using its API default because it had no effort parameter, while the others ran at xhigh.
Thoroughness did not guarantee the best result
Opus 4.8 offers a cautionary case. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem.
That outcome challenges a common assumption about AI: more analysis is not necessarily the same as better management. A useful agent must combine investigation, execution and respect for boundaries. Firmulate’s do-nothing baseline scored 26 because partial progress counts, but a single breach of trust caps the total. The governing principle is explicit: “no amount of good work outweighs a breach of trust.”
A company designed to make behavior visible
The live company contains 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, operates with a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday and decision is versioned and auditable, making the experiment watchable rather than a retrospective claim.
Readers can also test their instincts against 242 real, unedited management decisions in Firmulate’s “guess the model” quiz. The exercise reveals how difficult it can be to identify a model from confident business prose alone. What separates participants becomes clearer through their actions: whether they investigate, finish, escalate and preserve trust.

Test the difficult day before deployment
For travel and outdoor businesses, AI agents may eventually encounter customer records, booking problems, operational forecasts or urgent messages from managers. The Firmulate result is reassuring: all five participants resisted every manipulation attempt. It is also a reminder that refusal is only one dimension of readiness. Some models protected the company yet still failed to complete valuable, authorized work.
Enterprises can run the same kind of wargame against a read-only export of their own business, with nothing writing back to real systems. That creates an opportunity to examine behavior before an agent reaches live operations. The practical question is no longer simply whether an AI can produce a polished answer. It is whether it can remain honest, use the evidence available to it and complete the journey when pressure rises.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html