
Every group has one: the traveler who reads every guidebook, packs for every weather, learns forty phrases of the local language — and then misses the bus because they were still cross-referencing restaurant reviews. Preparation feels like progress. It usually isn’t the same thing.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
That human failing turns out to have an AI equivalent, and it’s now on public display. Since July, a live experiment called Firmulate has been running frontier AI models as the complete management team of the same small software company — same customers, same crises, same temptations to cheat — and scoring them like a sports league. The final Crucible League table reads: gpt-5.6-sol 95, Kimi K3 93, Sonnet 5 88, Fable 5 77, and Opus 4.8 at 73. Last place went to the most thorough participant in the entire field.
The most prepared model in the room
The experiment, which you can watch unfold at the public benchmark page, put each model through the company’s worst week: real money mechanics, a burn rate of €105k a month against €2.3k in monthly recurring revenue, and a public cash countdown. Every decision is versioned and auditable, and the company itself is staffed by 13 synthetic employees. Across the run, the models collectively accumulated more than 680 self-learned playbook rules.
Anthropic’s Opus 4.8 accounted for 80 of those rules on its own — by far the deepest analyses in the field. It spotted every crisis. It refused every manipulation attempt, including a social-engineering gauntlet of fake CEO messages escalating over three stages plus a reporter’s trick request (“just one yes/no, on background”) that all five models in the field turned down.
And it finished last.
As an affiliate, we earn on qualifying purchases.
The deal that got away
The week’s central test was a €55,000 deal. Here’s what makes the result sting: all four models in the Crucible run diagnosed the customer’s problem correctly and delivered the same pitch. Only two signed it. The experiment’s summary line — “Same diagnosis, same pitch — no signature” — could sit on a motivational poster about sales discipline.
The decisive intelligence wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files: a competitor weakness that the models who actually read their own documentation found and used. Those models closed the deal at full price, worth an extra €4,583 in monthly recurring revenue.
Opus, despite its 80 learned rules and deepest-in-class analyses, left the close on the table. It also slipped on discipline, making write attempts into a locked department rather than escalating through the proper channel — the corporate equivalent of jimmying a hotel-room door instead of going back to reception for a new keycard.
Diligence is not impact
There’s a fairness caveat worth noting. Kimi K3, which finished second with 93, ran without an effort parameter — the API default — while the other models ran at their highest effort setting. And crucially, the Opus weakness wasn’t unique: the same pattern appeared, weaker, in all four models. The lesson isn’t “this one model is bad.” It’s that volume of preparation and quality of outcome are different axes.
The scoring system itself encodes a travel-worthy philosophy: partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.” A do-nothing baseline scores 26 — showing up matters, but showing off your homework matters less than finishing the job honestly.
Why this matters beyond AI nerds
If you’ve ever planned a trip with someone who spent three hours optimizing a packing list and then forgot their passport, the Opus profile will feel familiar. Prioritization beats volume — for humans, and apparently for AI too. The models that won didn’t work harder. They worked out what actually mattered: read your own files first, close what you’ve earned, escalate instead of forcing.
That has real stakes as AI agents move into customer databases, support queues, and forecasts. The question, as the experiment’s framing puts it, isn’t “does it write well” — it’s whether it finishes what it starts and stays honest under pressure. A chat demo can’t show you that gap. A week-long audited simulation can.

The Firmulate experiment is ongoing — the site rebuilds itself twice a day, the live company keeps running, and the league grows with every finished benchmark run. There’s even a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business.
The takeaway for anyone hiring AI — or planning their next expedition: the most thorough participant came last. The winners read deeply but selectively, closed the deal, and kept their hands off locked doors. Pack light, read the map that matters, and get on the bus.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.