AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Every traveler knows the type. You hand two guides the same dense trip dossier — visa rules buried in an annex, a border-crossing restriction tucked into a footnote of a footnote — and one comes back with a flawless route while the other waves you onto a closed road. The difference isn’t intelligence. It’s whether they actually read the file before answering.

It turns out the same distinction now separates winning AI agents from losing ones — and for the first time, someone has measured it with real money on the line. Firmulate’s benchmark league, finalized in July 2026, ran four frontier AI models through the identical worst week of a small software company. The decisive fact in that week wasn’t shouted by a customer or hidden in a crisis. It was buried two document references deep in the company’s own files. The models that dug it up closed a €55,000 deal at full price. The ones that didn’t lost it automatically.

Same storm, four captains

The setup is elegantly cruel. Each frontier model got the same small software company, the same customers, the same crises, the same temptations to cheat — only the model changed. Every decision was versioned and auditable, so nothing could be quietly smoothed over afterward.

The crucible league’s final standings tell a sharp story:

  • gpt-5.6-sol — 95 points, described as “the complete performance”
  • Kimi K3 — 93, the Moonshot newcomer with the cleanest discipline in the field
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73, dead last despite being the most thorough participant

For context, a do-nothing baseline scores 26 — partial progress counts, but the scoring philosophy is unforgiving on one point: a single breach of trust caps the total. No amount of good work outweighs a breach of trust. It’s the expedition-leader standard: one act that endangers the group erases the whole trek.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

All four saw the weather. Two read the map.

Here’s the finding that should reset how anyone buys an AI agent. All four models spotted every crisis. All four refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

Why? The deal turned on a competitor weakness that wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. Models that followed the trail won the deal at full price — worth an additional €4,583 in monthly recurring revenue. Models that stopped one reference short never knew what they were missing.

It’s the AI equivalent of a guide who reads the trip report mentioning that the pass closes at noon, and one who reads only the summary. Both are fluent. Both are confident. Only one gets you over the mountain.

The pressure test worked

The experiment also staged social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five model runs refused. Kimi K3’s on-record reasoning stood out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.”

So honesty under pressure is now table stakes at the frontier. Finishing the job — reading the files, closing the loop — is where the field separates.

Thoroughness isn’t the same as follow-through

The most counterintuitive profile belongs to Opus 4.8: the most thorough participant, with 80 additional learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. A guide who researches every trail but never books the hut has planned a trip, not taken one.

You can watch the company run

None of this is a one-off lab report. Firmulate operates a live synthetic company — 13 employees, real money mechanics, burning €105,000 a month against €2,300 in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. It’s watchable in real time, like a base-camp feed for an expedition you can audit.

For those who want skin in the game, 242 real, unedited management decisions power a “guess the model” quiz — a surprisingly humbling parlor game. And enterprises can run the same wargame against a read-only export of their own business; nothing ever writes back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

One fairness note: Kimi K3 ran at its API-default effort setting while competitors ran at xhigh — and still took second. That detail cuts both ways, but the headline finding is robust: “reads your files before answering” is not a chat-demo nicety. It is a measurable, purchase-deciding property of AI agents, worth €55,000 in a single week in this simulation. Before you hire an AI workforce, wargame it. Ask not how well it writes, but whether it finishes what it starts — and whether it did the reading.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Customer Retention Strategies: Keeping Your Clients Loyal

Optimize your customer retention strategies to build lasting loyalty—discover proven methods that can transform your client relationships today.

Crisis Management 101: Preparing Your Business for the Unexpected

Harness essential strategies in Crisis Management 101 to prepare your business for the unexpected—discover how to stay resilient when it matters most.

The Standing Desk Mistake That Makes Posture Worse, Not Better

Better posture starts with proper desk height—discover the common mistake that can worsen your posture and how to avoid it.

Choice Hotels International Surges In Global Coverage

Choice Hotels International experiences a surge in worldwide media mentions, reflecting increased global attention on the company’s recent developments.