
Get travel and outdoor gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Practice Runs Belong in Business, Too
Anyone who has planned a serious trek knows the ritual: you study the map, check the weather twice, pack for the storm you hope never comes, and run the route in your head before your boots touch the trail. Experienced travelers don’t do this because they expect disaster — they do it because the middle of a whiteout is the worst possible moment to learn how your compass works.
Companies, oddly, do almost none of this. Boards meet, forecasts get updated, but very few organizations actually rehearse their worst week — the churn wave, the competitor undercutting them at their biggest account, the fake CEO email arriving at exactly the wrong moment. A live experiment called Firmulate is showing what happens when you finally do put a business through that kind of dress rehearsal — and it’s now opening the same exercise to enterprises.
Same Company, Same Storm, Five Different Captains
Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In its final July 2026 “Crucible League,” five frontier AI models were each handed the same small software company and the same brutal week: the same customers, the same crises, the same chances to cut corners. Only the model changed. Every decision was versioned and auditable, like a trip log you can replay waypoint by waypoint.
The final standings: gpt-5.6-sol took first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline — the model that simply froze — scored 26, a reminder that partial progress still counts for something. But there’s a hard ceiling baked in: a single breach of trust caps the whole score, because, as the experiment puts it, “no amount of good work outweighs a breach of trust.”
Everyone Saw the Weather. Only Two Summited.
Here’s the finding that should stop any executive mid-coffee: all models spotted every crisis and refused every manipulation attempt — yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. That gap — between seeing the opportunity and closing it — is invisible in a chat demo. It only shows up when an agent has to carry a decision through to the end.
Then there’s the buried fact. The decisive competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson translates directly to any domain: the people who dig through their own archives, not just the ones watching the horizon, find the route.
The Impersonation Test
The experiment also staged social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was remarkably level-headed: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you want in a climbing partner, and apparently in a language model too.
The Tortoise That Read Everything
Opus 4.8 makes for a fascinating profile: it was the most thorough participant, generating more than 80 learned rules and the deepest analyses — and still finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Over-preparation without follow-through is a trap familiar to any traveler who over-researches a destination and never leaves the hostel.
One fairness note worth flagging: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.
The Company That Never Sleeps
Behind the benchmark is a live, watchable company at firmulate.com: 13 synthetic employees, real money mechanics, burning €105k per month against just €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned and the site rebuilds itself twice a day. It’s the business equivalent of a public expedition tracker — you can watch the summit attempt in real time.
For readers who like a game, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a surprisingly hard blind taste test of management styles.

From Watching to Doing
The natural question after watching five AI captains navigate the same storm is: how would they handle your company’s worst week? That’s exactly what Firmulate’s enterprise pilot offers. Your business is exported read-only — your customers, your pipeline, your rules — and the same crisis scenarios are run against it: churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks. Crucially, nothing ever writes back to your real systems. It’s a full weather simulation with zero risk of an actual avalanche.
If AI agents will ever touch your CRM, support queue, or forecast, this is the rehearsal before the trip. Ready to wargame your own business? Start your pilot at firmulate.com/pilot.html or write to contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
