AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

AI tools sound polished. Can they actually run a business?

For readers evaluating automation, the familiar demonstrations leave a crucial gap. A model may summarize a meeting, draft a sales email or produce a confident strategy. None of that shows whether it will finish a difficult job, search for overlooked evidence or resist pressure to misuse its authority.

Firmulate turns that gap into an unusually revealing interactive article. Its guess-the-model quiz presents 242 real, unedited management decisions and asks readers to identify which frontier AI made each one. The exercise is entertaining, but the differences it exposes are consequential: the models faced identical business conditions, yet developed recognizable patterns in how they investigated, communicated and followed through.

Amazon

AI decision-making evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same company, the same terrible week

Firmulate placed each frontier model in charge of the same small software company during its worst week. The customers, crises and temptations were held constant. Every decision was versioned and auditable, allowing performance to be compared as management behavior rather than conversational style.

The live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, while a public cash countdown makes delay visible. Its workforce has accumulated more than 680 self-learned playbook rules, and every workday is versioned. The experiment is real, ongoing and watchable.

The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the evaluation imposed a hard ethical boundary: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

Shared awareness did not produce shared results

All models detected every crisis and rejected every manipulation attempt. That uniform success might suggest they were functionally interchangeable. The sales outcome showed otherwise. Only two signed the €55,000 deal their own analysis had earned. The central contradiction was stark: “Same diagnosis, same pitch — no signature.”

The difference was not hidden in a dramatic customer message. A decisive weakness in the competitor’s position sat two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The episode makes a practical distinction between an AI that understands the visible problem and one that searches the organization’s records before acting.

For companies considering agents for sales, support or operations, that distinction matters. An assistant can recognize a commercial opportunity and produce a persuasive recommendation while still failing to complete the action that creates value. The quiz makes this failure of follow-through easier to see because readers encounter the decisions directly, without a polished retrospective smoothing away hesitation or omission.

Pressure revealed a firmer consensus

The security test unfolded through fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3’s on-record reasoning was concise and operational: “Treat the request as a suspected approval-bypass / possible impersonation.”

This result is important precisely because the models differed elsewhere. Their management styles could diverge in depth, brevity and willingness to pursue information, but they converged when manipulation threatened authorization or confidentiality. In this field, personality did not mean abandoning basic trust boundaries.

Thoroughness was not the same as effectiveness

Opus 4.8 offers the clearest caution against equating volume with managerial quality. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The sales close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker form of that same problem appeared in all four of the other models.

The comparison does require one fairness note. Kimi K3 ran without an effort parameter, using its API default, while the others ran at xhigh. That condition does not erase its result, but it belongs beside any interpretation of the standings.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

business AI simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A more useful way to judge AI workers

Firmulate’s quiz works because it replaces branding and benchmark abstraction with recognizable workplace choices. Readers can compare a dissertation-like analysis, a terse escalation or a refusal to engage with distracting communication, then decide which model they trust to manage the next step.

The larger lesson is not that one communication style always wins. It is that management quality emerges across a chain of behavior: noticing a crisis, reading the relevant files, maintaining trust, escalating correctly and completing the commercially valuable action.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That turns the public experiment’s central question into a procurement test: before hiring an AI workforce, watch how it behaves during the worst week you can give it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management decision testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI performance assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Claude for Legal

Anthropic introduces ‘Claude for Legal,’ a suite of AI agents designed to streamline legal workflows across practice areas, with initial tools for contract review and compliance.

AI Users Rally Against Watermarks: Will They Limit Claude’s Use In Professional And Educational Settings?

Anthropic introduces machine-readable watermarks in Claude AI outputs, sparking protests from users concerned about detection in workplaces and schools.

Fable 5 Vs. GPT-5.6 Sol On An NP-Hard Problem: Does /Goal Help?

Researchers compare Fable 5 and GPT-5.6 Sol on an NP-hard problem, analyzing whether the /goal prompt improves solution accuracy.

Show HN: TikTok but for Scientific Papers

Papel launches as a TikTok-like platform for academic research, offering personalized feeds, AI insights, and community features for scientists and students.