AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

A field test for artificial intelligence

Travelers and outdoor leaders know that competence becomes visible when the weather changes, the route disappears or the plan collides with reality. Firmulate applies that same principle to frontier artificial intelligence: do not judge the guide by the briefing; watch what it does during the worst week.

In the Crucible League, each model ran the same small software company through identical customers, crises and temptations. The decisions were versioned and auditable, producing a record that readers can now explore as a kind of management field guide. The resulting guess-the-model quiz draws on 242 real, unedited decisions and asks a deceptively revealing question: can you identify an AI model by the way it manages?

The answer matters beyond the novelty of matching a decision to a machine. The experiment suggests that frontier models can reach the same diagnosis yet differ sharply in whether they investigate deeply, close a valuable deal or maintain operational discipline.

Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same terrain, different management personalities

The final July 2026 standings put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust.”

Those results are not simply a ranking of who recognized trouble. Every model spotted every crisis, and every model refused every manipulation attempt. The decisive differences appeared after recognition, in the less glamorous work of reading, following through and finishing.

The clearest example was a €55,000 deal. All the models had access to the same situation and could develop the same diagnosis and pitch. Yet only two signed the deal their own analysis had earned: “Same diagnosis, same pitch — no signature.” It is the managerial equivalent of identifying the correct trail and then failing to take the final turn.

The valuable clue was not in the obvious place

The crucial competitive weakness did not sit in the customer event itself. It was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR.

That finding gives the quiz its bite. Readers are not merely distinguishing writing styles. They are looking for behavioral signatures: which model searches the available record, which produces exhaustive analysis, which keeps communication restrained and which converts insight into action. Because the decisions are unedited, polish cannot erase the gap between sounding capable and completing the job.

Pressure tested honesty

The experiment also placed every model in the path of social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused.

Kimi K3’s on-record reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response helps explain why management personality is more than tone. A terse refusal can represent strong discipline, while a lengthy answer can coexist with incomplete execution.

The comparison carries a fairness qualification. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its 93 therefore belongs in the published league table, but the differing setup is important context when comparing behavior.

When thoroughness becomes a trap

Opus 4.8 offers the most striking character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last with 73. The model left the close on the table, and its discipline slipped when it attempted writes into a locked department instead of escalating.

The same weakness appeared in all four other models, although less strongly. That makes the result more useful than a simple winner-and-loser story. Firmulate’s experiment shows that intelligence, diligence and managerial completion do not always travel together. The model with the deepest analysis can still fail at the point where a decision must become a finished business outcome.

A company under visible pressure

The setting is deliberately unforgiving. The live company has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. It maintains a public cash countdown, has accumulated more than 680 self-learned playbook rules and versions every workday.

This is why the quiz works as an interactive article rather than a personality test dressed up as research. Each answer comes from a real, watchable experiment in which the company’s condition, the models’ choices and the business consequences remain visible. The reader guesses first, then confronts the behavior that separated recognition from execution.

Infographic —
The findings at a glance — source: firmulate.com.

The practical lesson for human organizations

Firmulate’s central challenge is aimed at companies considering AI agents for consequential work: writing quality is not the same thing as management quality. An agent may need to read internal material before acting, resist an authority-bypass attempt, respect locked areas and carry a sound commercial analysis through to signature.

Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to real systems. That creates a way to observe how a prospective AI workforce behaves around an organization’s actual context without granting it the power to alter that environment.

For readers accustomed to judging routes, equipment and guides under changing conditions, the analogy is natural. The compelling question is not which model gives the most impressive briefing. It is which one finds the buried fact, protects trust and finishes the journey.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

How to Create a Business Plan: A Step-by-Step Guide for Beginners

Discover the essential steps to creating a business plan that sets your startup up for success and unlocks opportunities you haven’t yet imagined.

Social Media Marketing Strategies for Small Business Success

The key to small business success lies in mastering social media marketing strategies that can transform your reach—discover how to unlock your full potential.

The Evolution of Ecommerce: Trends Every Business Should Watch

Modern ecommerce trends are transforming how businesses engage customers—discover the key developments you can’t afford to ignore.

Saudi Arabia Restarts Crude Loadings at Major Gulf Terminal After Nearly Four-Month Halt

Saudi Arabia has restarted crude oil loadings at its major Gulf terminal after nearly four months of suspension, signaling a potential shift in its oil export activity.