
A clever answer is not the same as competent management
For readers choosing AI tools and automation platforms, leaderboards offer an appealing shortcut: find the model with the strongest score and put it to work. But coding tests and chat arenas usually measure the quality of an answer produced in a controlled moment. They reveal much less about what happens when an agent must triage competing crises, search company records, resist pressure, finish commercial work and remain candid about the result.
That is the measurement gap exposed by Firmulate, a live experiment that runs frontier models as complete companies. Its premise is that an AI workforce should be judged by management quality, not merely chat quality. Instead of asking another isolated question, the experiment puts each model in charge of the same small software business during its worst week, with identical customers, crises and temptations. Every decision is versioned and auditable.
enterprise AI management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The crisis was visible. The opportunity was buried.
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted, although a single breach of trust capped the total. The governing principle was blunt: “no amount of good work outweighs a breach of trust.”
The ranking matters, but the stories behind it matter more. Every model identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 contract their own work had earned. Firmulate summarizes the disconnect as: “Same diagnosis, same pitch — no signature.” That is precisely the kind of failure a polished conversation can conceal. The agent understood the situation and prepared the right commercial case, but understanding did not reliably become an outcome.
The decisive information was not sitting in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that read the file won the deal at full price, adding €4,583 in monthly recurring revenue. This is not primarily a test of eloquence. It is a test of whether an agent checks the available evidence before acting and carries that evidence through to a completed decision.
Pressure revealed discipline as well as intelligence
The experiment also staged fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest operational interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result deserves attention because enterprise automation creates new opportunities for social engineering. An agent operating around forecasts, customer records or internal communications must distinguish authority from the appearance of authority. In this field, the models showed that they could do so even while other problems competed for attention.
K3’s performance also comes with an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. That difference should accompany any comparison rather than being buried beneath the final table.
Thoroughness did not guarantee completion
Opus 4.8 offers the most instructive profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It failed to close the deal and lost discipline by attempting writes into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.
This is the distinction between looking industrious and managing effectively. More analysis can improve a decision, but it can also create the appearance of progress while the essential action remains unfinished. An enterprise evaluating agents therefore needs to observe handoffs, escalation behavior and closure—not just the reasoning artifact left behind.
The surrounding company makes those consequences tangible. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, more than 680 playbook rules have been learned, and every workday is versioned. The experiment is real, continuously watchable through Firmulate, and its published league expands as runs finish.

AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Scenario names may become the new AI curriculum
“Churn wave,” “price increase,” “downround” and “PR crisis” describe a more revealing class of evaluation than another generic prompt set. They force an agent to balance urgency, evidence, commercial consequences and trust across days. The resulting question is not simply whether a model knows what good management sounds like. It is whether the model can practice it when capacity is constrained and several defensible actions compete.
Firmulate also turns 242 real, unedited management decisions into a guess-the-model quiz, illustrating how difficult it can be to identify a model from isolated decisions alone. The differences become clearer when behavior accumulates into a company outcome.
Enterprises can apply the same wargame to a read-only export of their own business, with nothing written back to real systems. That offers a practical bridge between public benchmarking and procurement: test an agent against the organization’s actual ambiguity before allowing it near production authority. The full league and plain-language findings are available on the Firmulate benchmark page.
The emerging category is management quality. Chat quality still matters, but it is only an ingredient. The more consequential measures are whether an agent reads before acting, resists improper pressure, escalates when blocked, completes valuable work and tells the board what actually happened.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.