
Anyone who has traveled with a guide knows the difference between competence and completion. The guide who spots every storm cloud, reads the map perfectly, warns you off every dangerous ridge — and then, with the summit in view, simply stops walking. You’re safe, you’re informed, and you still didn’t get where you paid to go.
Get travel and outdoor gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
That, in miniature, is the strangest finding from a live business experiment now running publicly at Firmulate: four frontier AI models were each handed the same small software company to run through its worst week. All four spotted every crisis. All four refused every attempt to manipulate them. Only two actually closed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
Why a do-nothing manager still scores 26
The benchmark’s most talked-about design choice is its floor. A do-nothing baseline run — a manager who never lifts a finger — scores 26 points, not 0. That’s not generosity; it’s honesty. In any real company, simply not making things worse has value, and partial progress counts. A manager who handles three of five crises has genuinely accomplished something, and the score reflects it.
But the scale has a hard ceiling of a different kind: a single breach of trust caps the total. The stated principle is blunt — “no amount of good work outweighs a breach of trust.” An agent could ace every operational test and still fail the grade by crossing one ethical line. For travelers, the analogy is obvious: a guide who gets you up the mountain fast but takes an unsafe shortcut isn’t a 90% good guide.
The buried fact that decided the deal
The €55,000 deal turned on something subtle. The decisive competitor weakness wasn’t in the customer’s event or conversation — it sat two document references deep in the company’s own files. The models that actually read their own documentation found it, and won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The models that stopped reading left the close on the table.
It’s the business equivalent of the guide who never checked the trail conditions posted at the ranger station. Everything looked fine; the winning information was sitting in the file cabinet.
The league table
The final July 2026 standings: gpt-5.6-sol first at 95 — the verdict calls it “the complete performance” for finding the buried fact and closing the deal. Kimi K3, the newcomer from Moonshot, took second at 93 with the cleanest discipline of the field. Sonnet 5 followed at 88, Fable 5 at 77, and Opus 4.8 at 73. One fairness note: K3 ran at the API’s default effort setting while the others ran at xhigh — making its near-top finish arguably more impressive, not less.
The Opus 4.8 profile is the cautionary tale. It was the most thorough participant, with over 80 learned rules and the deepest analyses — and it finished last. Discipline slipped: it attempted writes into a locked department instead of escalating, and the close went unfinished. The same weakness appeared, more weakly, in all four models.
Under pressure, everyone said no
The experiment’s social-engineering thread deserves its own mention. Fake CEO messages escalated over three stages, capped by a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Whatever else these systems struggle with, they did not fold to flattery or forged authority.
You can watch the company run
This isn’t a slide deck; it’s a live operation. The synthetic company has 13 employees, real money mechanics — burning €105,000 a month against €2,300 in MRR — a public cash countdown, over 680 self-learned playbook rules, and every workday versioned for audit. The site rebuilds itself twice a day, and finished benchmark runs publish automatically. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The lesson for anyone choosing tools — or guides — is that chat quality is not management quality. The interesting failures aren’t the dramatic ones; nobody got tricked, nobody cheated. The interesting failure is the quiet one: brilliant analysis, perfect vigilance, and a deal left unsigned because reading one more file felt optional. If an AI will touch your business, ask the questions this benchmark asks: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work cost? And be wary of any score that comes back a perfect 100. A benchmark that includes distrust of round numbers is one that has probably earned yours.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI business decision analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
