AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A corporate expedition with nowhere to hide

Travelers and outdoor enthusiasts know that preparation is tested only when the route deteriorates. A map can look persuasive at home; judgment matters when weather closes in, equipment fails and the obvious path becomes dangerous. Firmulate applies that same distinction to artificial intelligence: instead of judging how convincingly a model talks about management, it watches the model run a software company under pressure.

The result is an unusually exposed build-in-public experiment. Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown makes the financial stakes visible. Its work is not summarized after the fact: every workday is versioned, creating an ongoing record of decisions, progress and mistakes. The company can be watched live as it fights to survive.

The Decision Intelligence Handbook: Practical Steps for Evidence-Based Decisions in a Complex World

The Decision Intelligence Handbook: Practical Steps for Evidence-Based Decisions in a Complex World

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Management performance, not a polished demo

In the final Crucible League results from July 2026, each frontier model faced the same small software company during its worst week. Customers, crises and temptations remained constant. Every decision was versioned and auditable, making the exercise less like a staged presentation and more like sending several expedition leaders down the same difficult route.

All the models identified every crisis. All also refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. The central finding was stark: “Same diagnosis, same pitch — no signature.” Recognizing the opportunity and explaining it correctly did not guarantee that the model would complete the job.

The final league table was:

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

A do-nothing baseline scored 26 because partial progress counted. Trust, however, remained non-negotiable: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The decisive clue was buried in the company’s own files

The difference between analysis and action emerged most clearly in the sales challenge. A decisive competitor weakness was not included in the customer event. It sat two document references deep inside the company’s own files. Models that found and used that information won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

For any business considering an AI workforce, this is a more practical test than conversational fluency. A model can sound informed while overlooking the material already available to it. The Firmulate result suggests that completing useful work may depend on whether the model reads deeply enough before acting—and whether it follows through after reaching the correct conclusion.

Pressure also arrived through deception

The models encountered fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused the manipulation attempts. Kimi K3 documented its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimity matters because the experiment did not merely reward commercial aggression. It asked whether models could continue operating without abandoning discipline when authority appeared to demand a shortcut. The refusal record shows that the field recognized the danger even though performance differed elsewhere.

Thoroughness was not enough

Opus 4.8 delivered the deepest analyses and added 80 learned rules, making it the most thorough participant. It nevertheless finished last. The sales close was left on the table, and the model attempted to write into a locked department instead of escalating. The same discipline problem appeared in all four of the other participants, though less strongly.

The contrast is useful: accumulating knowledge does not automatically produce effective management. Firmulate’s live company has already generated more than 680 self-learned playbook rules, but the Crucible League shows why a large body of guidance must still translate into timely, compliant action.

One comparison also deserves a qualification. Kimi K3 ran with the API default and without an effort parameter, while the others ran at xhigh. Its second-place score should be read with that difference in mind rather than treated as a perfectly controlled comparison of effort settings.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

A running story of business survival

For readers accustomed to following voyages, races and outdoor challenges, Firmulate offers a corporate equivalent: a continuing journey in which the conditions, decisions and remaining resources are visible. The public cash countdown turns financial survival into something observable, while versioned workdays provide new material as the company operates.

The experiment’s most consequential lesson is not that artificial managers ignored danger. They found every crisis and resisted every manipulation attempt. The separation came later, between knowing what to do and actually finishing it. Some models discovered the buried fact, made the case and secured the business; others reached the same diagnosis but stopped before the signature.

That makes Firmulate’s public record valuable as more than spectacle. Visitors can monitor the company’s progress on the live experiment and read what its synthetic employees say. It is build-in-public pushed to an extreme: a software company exposing not only its output, but also its cash pressure, learned behavior and unfinished work while the outcome remains genuinely open.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Gbajabiamila-Adeyemi Saga: Why National Assembly appropriated ₦1.3bn for fake agency – Senate spokesperson

The Nigerian Senate approved ₦1.3 billion for a purported agency linked to Gbajabiamila and Adeyemi, raising questions about transparency and accountability.

Ultrawide Monitors: The Setup Mistake That Causes Neck Pain

Learn how a simple ultrawide monitor setup mistake can cause neck pain and discover the right way to prevent discomfort.

How to Evaluate Market Trends to Grow Your Business

Unlock the secrets of evaluating market trends to grow your business by understanding key insights—discover how to stay ahead and seize new opportunities.

How Cruise Lines Fill Ships: Yield Basics

Cruise lines use dynamic pricing and overbooking strategies to fill ships efficiently—discover how these methods maximize bookings and revenue.