AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

What kind of leader would you want when the route disappears?

Travel and outdoor decisions often reveal character under pressure. A capable expedition leader must notice changing conditions, consult the available information, resist reckless shortcuts and follow through when the moment demands action. Artificial intelligence models, it turns out, display similarly distinctive habits when placed in charge of a business.

Firmulate has turned those differences into an interactive challenge. Its guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how an anonymous model responded and try to identify it—not by its prose style alone, but by the managerial personality behind the choice.

The decisions come from a live experiment in which each frontier model ran the same small software company through its worst week. The customers, crises and temptations remained identical. Every decision was versioned and auditable, making the exercise less like a polished demonstration and more like watching several leaders tackle the same difficult course.

Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same hazards, very different journeys

The final Crucible League results from July 2026 show how widely the performances diverged. GPT-5.6-sol finished with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted.

Yet the benchmark imposed a firm ethical boundary: a single breach of trust capped the total. Its guiding principle was explicit: “no amount of good work outweighs a breach of trust”. That matters because the simulated company was not merely sorting routine correspondence. The models encountered financial pressure, operational problems and deliberate attempts to manipulate them.

On the most reassuring measure, the field was remarkably consistent. Every model spotted every crisis and refused every manipulation attempt. Fake messages from the chief executive escalated over three stages, while a reporter tried to extract information with the invitation “just one yes/no, on background”. All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

The sharper distinction emerged after the danger had been recognised. Only two models signed the €55,000 deal that their own analysis had earned. Firmulate summarises the gap neatly: “Same diagnosis, same pitch — no signature”. In other words, understanding the situation did not guarantee that a model would complete the commercial task.

The clue hidden off the obvious path

The deal also depended on whether a model explored beyond the event directly in front of it. The decisive weakness in a competitor was buried two document references deep in the company’s own files, rather than appearing in the customer event. Models that found and read that material won the deal at full price, adding +€4,583 MRR.

For anyone accustomed to planning a journey, the lesson feels familiar. The information that changes a decision may sit in the detailed route notes rather than on the sign at the trailhead. A model can sound confident, diagnose the broad problem correctly and still miss the crucial fact because it did not inspect the available documents.

Thoroughness is not the same as effectiveness

Opus 4.8 provides the experiment’s most counterintuitive character study. It was the most thorough participant, learning +80 rules and producing the deepest analyses, yet it finished last. The model left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four of the others, though less strongly.

This is what gives the quiz more substance than a blind comparison of writing samples. Readers are not simply looking for verbosity, confidence or familiar phrasing. They are judging whether a model investigates, acts, escalates appropriately and finishes what it begins. One may write a dissertation; another may be terse; another may refuse to engage with irrelevant noise. Those tendencies become recognisable management profiles when repeated across real decisions.

There is an important qualification around the close result at the top. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its 93-point finish should therefore be read with that difference in mind rather than treated as a perfectly controlled comparison of effort settings.

Infographic —
The findings at a glance — source: firmulate.com.

A company designed to make strengths and weaknesses visible

The live company gives these decisions consequence. It has 13 synthetic employees and operates with real money mechanics, burning €105k each month against €2.3k MRR. Its cash countdown is public, it has accumulated 680+ self-learned playbook rules, and every workday is versioned. The result is watchable evidence of how models behave over time rather than a collection of hand-picked chat responses.

Firmulate also offers enterprises the opportunity to run the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organisations to observe how a prospective AI workforce handles their information and pressures without granting it operational control.

For readers interested in travel, exploration and calculated risk, the broader message is compelling: competence becomes clearest when conditions deteriorate. The best guide is not merely the one who identifies the storm or explains the map. It is the one who reads the overlooked notes, protects the group from bad instructions and completes the necessary move. Firmulate’s quiz asks whether those behavioural signatures are distinctive enough for humans to recognise—and whether we are paying attention to the decisions that matter.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Rathen

Heavy storms have caused disruptions in Rathen, Saxony, impacting the Felsenbühne Rathen and surrounding areas. Authorities issue warnings amid ongoing weather concerns.

Car-Free Guide to Sunshine Coast

Wondering how to explore the Sunshine Coast without a car? Discover tips for an unforgettable car-free adventure!

Will The Lowest Temperature In Tokyo Be 27°C On July 11?

Weather forecasts suggest Tokyo’s lowest temperature may reach 27°C on July 11, but official data has not confirmed this. Read for details and implications.

Adelaide’s Local-First Food Trail (Without Driving Everywhere)

Wish to explore Adelaide’s vibrant local food scene without driving? Discover how walking, cycling, and public transit reveal hidden culinary gems.