AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Every traveler knows the difference between a guidebook and a good guide. The guidebook answers the question you asked. The guide notices the storm rolling in, reads the fine print you skipped, and gets you out before the airport closes.

We keep hearing that AI has become the ultimate guidebook — models top coding leaderboards, ace chat arenas, and draft itineraries that read like they were written by a Condé Nast veteran. But a company called Firmulate is asking a different question: what happens when the AI isn’t answering questions anymore, but actually running the operation during its worst week? The answer, from a recently completed experiment, is humbling for the entire industry.

Same storm, four captains

Firmulate’s premise is simple and slightly brutal. Four frontier AI models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so there’s no arguing about what happened.

Think of it as the corporate equivalent of dropping four hikers into the same canyon with the same gear and the same forecast — then watching who actually reads the map before the river rises.

The leaderboard

The final standings from the July 2026 league run tell the story:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As Firmulate puts it: no amount of good work outweighs a breach of trust.

The finding that chat demos can’t show

Here’s the part that should make anyone evaluating AI agents sit up. All the models spotted every crisis. All of them refused every manipulation attempt. And yet only two of them signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

The decisive detail was buried two document references deep in the company’s own files — a competitor weakness that had nothing to do with the customer event itself. The models that actually read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. It’s the AI equivalent of the traveler who checks the visa fine print instead of trusting the headline on the booking site.

Pressure and honesty

The week also included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

The thoroughness trap

The most striking profile belongs to Opus 4.8: the most thorough participant in the field, with 80 learned rules and the deepest analyses — and last place. The deal was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four models. One fairness note worth flagging: K3 ran at the API’s default effort setting while the others ran at maximum, making its second-place finish even more striking.

You can watch it lose money

None of this is a slide deck. Firmulate runs a live, synthetic company — 13 employees, real money mechanics, burning €105,000 a month against €2,300 in MRR, with a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned, and the whole thing is watchable at firmulate.com. There’s also a quiz built from 242 real, unedited management decisions, where you guess which model did what. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full benchmarks and plain-language findings are public.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The lesson travels well beyond software companies. Whether an AI is managing a support queue, a supply chain, or a portfolio of hotel partnerships, chat quality is not management quality. Closing the loop, reading the file, staying honest when a fake CEO comes calling — those are the skills that decide whether your worst week ends with a signed deal or a silent failure. The models that win leaderboards write beautiful answers. The models that survive a price war, a churn wave, or a PR crisis finish the job. Measure the second thing.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI travel planner app

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Independence Day: The Revolution That Chose Freedom Over Power

A historic revolution on July 4th led to the founding of a nation prioritizing liberty over authority, marking a pivotal shift in history.

Delta Airlines flight diverted to Fresno Yosemite International Airport

A Delta Airlines flight was diverted unexpectedly to Fresno Yosemite International Airport due to an emergency. Details are still emerging.

The Standing Desk Mistake That Makes Posture Worse, Not Better

Better posture starts with proper desk height—discover the common mistake that can worsen your posture and how to avoid it.

Document Scanners: How to Go Paperless Without Losing Your Mind

Losing paper clutter starts with the right scanner—discover essential tips to streamline your digital transition without the stress.