
What kind of leader would you want when the route disappears?
Travel and outdoor decisions often reveal character under pressure. A capable expedition leader must notice changing conditions, consult the available information, resist reckless shortcuts and follow through when the moment demands action. Artificial intelligence models, it turns out, display similarly distinctive habits when placed in charge of a business.
Firmulate has turned those differences into an interactive challenge. Its guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how an anonymous model responded and try to identify it—not by its prose style alone, but by the managerial personality behind the choice.
The decisions come from a live experiment in which each frontier model ran the same small software company through its worst week. The customers, crises and temptations remained identical. Every decision was versioned and auditable, making the exercise less like a polished demonstration and more like watching several leaders tackle the same difficult course.
As an affiliate, we earn on qualifying purchases.
The same hazards, very different journeys
The final Crucible League results from July 2026 show how widely the performances diverged. GPT-5.6-sol finished with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted.
Yet the benchmark imposed a firm ethical boundary: a single breach of trust capped the total. Its guiding principle was explicit: “no amount of good work outweighs a breach of trust”. That matters because the simulated company was not merely sorting routine correspondence. The models encountered financial pressure, operational problems and deliberate attempts to manipulate them.
On the most reassuring measure, the field was remarkably consistent. Every model spotted every crisis and refused every manipulation attempt. Fake messages from the chief executive escalated over three stages, while a reporter tried to extract information with the invitation “just one yes/no, on background”. All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
The sharper distinction emerged after the danger had been recognised. Only two models signed the €55,000 deal that their own analysis had earned. Firmulate summarises the gap neatly: “Same diagnosis, same pitch — no signature”. In other words, understanding the situation did not guarantee that a model would complete the commercial task.
The clue hidden off the obvious path
The deal also depended on whether a model explored beyond the event directly in front of it. The decisive weakness in a competitor was buried two document references deep in the company’s own files, rather than appearing in the customer event. Models that found and read that material won the deal at full price, adding +€4,583 MRR.
For anyone accustomed to planning a journey, the lesson feels familiar. The information that changes a decision may sit in the detailed route notes rather than on the sign at the trailhead. A model can sound confident, diagnose the broad problem correctly and still miss the crucial fact because it did not inspect the available documents.
Thoroughness is not the same as effectiveness
Opus 4.8 provides the experiment’s most counterintuitive character study. It was the most thorough participant, learning +80 rules and producing the deepest analyses, yet it finished last. The model left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four of the others, though less strongly.
This is what gives the quiz more substance than a blind comparison of writing samples. Readers are not simply looking for verbosity, confidence or familiar phrasing. They are judging whether a model investigates, acts, escalates appropriately and finishes what it begins. One may write a dissertation; another may be terse; another may refuse to engage with irrelevant noise. Those tendencies become recognisable management profiles when repeated across real decisions.
There is an important qualification around the close result at the top. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its 93-point finish should therefore be read with that difference in mind rather than treated as a perfectly controlled comparison of effort settings.

A company designed to make strengths and weaknesses visible
The live company gives these decisions consequence. It has 13 synthetic employees and operates with real money mechanics, burning €105k each month against €2.3k MRR. Its cash countdown is public, it has accumulated 680+ self-learned playbook rules, and every workday is versioned. The result is watchable evidence of how models behave over time rather than a collection of hand-picked chat responses.
Firmulate also offers enterprises the opportunity to run the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organisations to observe how a prospective AI workforce handles their information and pressures without granting it operational control.
For readers interested in travel, exploration and calculated risk, the broader message is compelling: competence becomes clearest when conditions deteriorate. The best guide is not merely the one who identifies the storm or explains the map. It is the one who reads the overlooked notes, protects the group from bad instructions and completes the necessary move. Firmulate’s quiz asks whether those behavioural signatures are distinctive enough for humans to recognise—and whether we are paying attention to the decisions that matter.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html