AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Every seasoned traveler knows the rule: never hire a guide on the strength of the brochure. You ask around, you check whether they’ve actually walked the route, and you watch how they behave when the weather turns. A polished pitch tells you nothing about what happens on day three, halfway up the mountain, when the plan falls apart.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get travel and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Businesses are now hiring AI models the way naive tourists hire guides — based on demos, reputation, and fluent conversation. A live experiment called Firmulate is trying to change that. It runs frontier AI models as the complete management team of a small software company through its worst week, with real money mechanics and versioned, auditable decisions — and its latest results contain a genuine surprise: Moonshot’s Kimi K3, the newcomer, finished second out of five, ahead of three of four Western frontier models.

The Worst Week in Business, Run Five Times

The setup is elegantly brutal. Each model got the same job: run the same small software company through identical crises, with the same customers, the same temptations to cut corners, and the same shot at a career-defining deal. Only the model changed. Every decision was versioned and auditable, so nothing rests on anecdotes.

The final July 2026 Crucible league table:

  • gpt-5.6-sol — 95 (found the buried fact, closed the deal)
  • Kimi K3 (Moonshot) — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own framing puts it: “no amount of good work outweighs a breach of trust.”

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Newcomer Actually Did

K3’s week reads like the résumé of a hire you’d want. It found the buried security needle — a decisive competitor weakness hidden two document references deep in the company’s own files, not in the customer event in front of it. Models that read the file won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. K3 closed it.

It also saved the churning customer and resisted every manipulation attempt thrown at it. The social engineering gauntlet included fake CEO messages escalating over three stages and a reporter’s disarming “just one yes/no, on background” trick. All five models refused all the baits — but K3’s on-record reasoning stood out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the whole week, K3 recorded just one deviation — the cleanest discipline in the field.

One fairness note belongs here, in the interest of honest scorekeeping: K3 ran without an effort parameter (the API default), while the other models ran at the maximum “xhigh” setting. Even so, the result stands as run.

The Finding That Should Worry Every Buyer

The strangest result wasn’t about K3 at all. Every model spotted every crisis and refused every manipulation. But only two of the five signed the €55,000 deal their own analysis had earned. The experiment’s summary is damning in its calm: “Same diagnosis, same pitch — no signature.”

That gap — between doing the analysis and finishing the job — is invisible in a chat demo. It only shows up when a model has to run something end to end, which is precisely how enterprises are starting to deploy AI into CRMs, support queues, and forecasts.

Then there’s Opus 4.8, the cautionary tale. It was the most thorough participant in the field: over 80 learned rules and the deepest analyses of anyone. It finished last. The deal was left on the table, and discipline slipped — write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Thoroughness, it turns out, is not the same thing as judgment.

Not a Slide Deck — A Company Losing Money in Public

Firmulate’s live company is real running software, not a simulation on paper: 13 synthetic employees, a burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it lose money in real time at firmulate.com — the site rebuilds itself twice a day.

For the curious, 242 real, unedited management decisions from the experiment power a “guess the model” quiz, and full benchmark results with plain-language findings are published openly. Enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The lesson travels well beyond software, and it’s the same one any backpacker learns at a trailhead: the flashiest credentials don’t predict performance under pressure — the route does. gpt-5.6-sol won this particular week, but the bigger story is that a newcomer from Moonshot beat three of four Western frontier models on management quality, not chat quality. The league is open.

If you’re choosing an AI model — or a guide — without running it through your own worst week, you’re not making a decision. You’re making a bet. The tools to test first now exist at Firmulate’s benchmarks. Check the map before you check the weather.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why NAS Systems Make Sense Only With a Real Backup Plan

Just relying on NAS alone isn’t enough; a comprehensive backup plan is essential to truly protect your data from unforeseen threats and failures.

Treadmill Desks: How to Walk Without Typing Like a Toddler

Learn how to walk comfortably at a treadmill desk without sacrificing productivity—discover tips to walk smoothly and work efficiently.

The Cruise Line Loyalty Question That Matters More Than Points

Inevitably, the most important loyalty question on a cruise isn’t about points but about how truly valued and satisfied you feel—discover why.

Chicago Plane Fireworks Strike

A plane at Chicago Midway Airport was hit by fireworks during landing, causing minor damage. Authorities are investigating the incident.