AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get travel and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What a Guide Who Did Nothing Would Score

Imagine hiring a mountain guide who never reads the route notes, never checks the weather, and freezes at every decision point — but at least never leads you off a cliff. Would you score them zero? Probably not. They kept you safe, which counts for something. But you wouldn’t hire them again either.

That, in miniature, is the puzzle behind one of the more unusual AI benchmarks published this year. Firmulate, a project that runs AI models as complete companies through simulated crises, gave its lowest-performing model 73 points out of 100 — and gave a “do-nothing” baseline run a full 26 points instead of zero. For travelers used to five-star review inflation and round-number scores, it’s a refreshing oddity. Here’s why the math works that way.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Four AI Models, One Terrible Week

Firmulate gave four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed, and every decision was versioned and auditable — like a summit log you can read line by line.

The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. The headline finding was striking in its symmetry: every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the researchers put it: “Same diagnosis, same pitch — no signature.”

Why Doing Nothing Scores 26, Not 0

The benchmark’s philosophy is that partial progress counts. A run that simply avoids disasters — no breaches of trust, no panicked missteps — has genuinely done something valuable, the way a cautious trekker who turns back early still got the group home safely. So the floor sits at 26 points rather than zero. A model has to actively accomplish things to climb above it.

But there’s a hard ceiling in the other direction: a single breach of trust caps the total grade, full stop. The benchmark’s stated principle is blunt — “no amount of good work outweighs a breach of trust.” You can think of it like a guide who is brilliant on the trail but once falsifies a safety log. For a traveler, that’s the person you never go out with again.

The Buried Fact That Won the Deal

The €55,000 deal didn’t hinge on anything the customer said. The decisive competitive weakness sat two document references deep in the company’s own files — the kind of detail only a model that actually read everything would find. The models that did read it won the deal at full price, worth +€4,583 in monthly recurring revenue. It’s the business equivalent of the guide who actually read the permit file and knew the pass closed at noon.

Thoroughness Isn’t Everything

The most instructive profile is Opus 4.8. It was the most thorough participant — over 80 learned rules, the deepest analyses — and still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating properly. The same weakness appeared, weaker, in all four models. Reading every map in the hut is not the same as getting the group over the pass.

One fairness note: Kimi K3 ran without an effort parameter while the others ran at maximum effort, and it still nearly topped the table.

Social Engineering, Refused Five for Five

The wargame included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five model runs refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

You Can Watch It Live

Unlike most AI research, this one is running in public. Firmulate operates a live synthetic company with 13 employees, real money mechanics — burning €105k per month against €2.3k in monthly recurring revenue — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com, where the site rebuilds itself twice a day and the league grows with every finished run. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Honest Benchmark’s Real Lesson

For anyone who relies on reviews — whether for a trekking operator or an AI vendor — Firmulate’s scoring is a small masterclass in honesty. Partial progress is worth something, so the floor isn’t zero. But trust is binary, so one breach caps everything. And a distrust of round 100s keeps the top of the table meaningful: even the best model left something on the table, finishing at 95 rather than a suspiciously perfect score.

The gap that matters — between diagnosing a problem and finishing the job — is invisible in chat demos. It only shows up when you watch an AI run something real, badly, under pressure, in public. That’s worth remembering the next time an AI product promises to handle your bookings, your itinerary, or your business. Ask not whether it writes well. Ask whether it finishes what it starts, reads your files first, and stays honest when nobody’s checking. The full league table and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Sustainable Neighborhoods in Sunshine Coast: Where to Explore

Explore the vibrant sustainable neighborhoods of Sunshine Coast, where eco-friendly living thrives and community connections blossom; discover the hidden gems waiting for you!

Hotspots

Just as hotspots reveal Earth’s dynamic changes, exploring their impact uncovers crucial insights into our environment and future risks.

Darwin Wet Season Reality Check: Travel Smart, Travel Light

Find out how to travel smart and light during Darwin’s wet season to stay dry and comfortable—before the weather catches you off guard.

Wildlife Viewing Etiquette in Freycinet National Park

Safeguard wildlife in Freycinet National Park with essential etiquette tips that ensure memorable encounters—discover how your actions can make a difference.