
Imagine an artificial intelligence not just generating text or answering questions, but actually running a tiny company — making real decisions, handling crises, and even losing money every day. This is not fiction. It’s the live experiment conducted by Firmulate, an open-door showcase of AI in management, where the stakes are high, and the outcomes are revealing.
The Live Company: A Window into AI’s Management Skills
At the core of this experiment is a small, simulated company staffed by 13 synthetic employees, working in a real-time environment that tracks its cash burn, customer crises, and decision-making processes. With a monthly cash burn of €105,000 against a modest €2,300 monthly recurring revenue, this company is clearly struggling, providing a challenging testbed for AI decision models.
Every day, this company faces the same set of crises—security threats, customer negotiations, and ethical dilemmas—and all decisions are recorded, versioned, and auditable. The goal? To see whether AI models can navigate these complexities honestly while pulling off profitable deals when possible.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experimental Setup: Four Top AI Models in Action
The experiment pits four frontier language models against each other:
- GPT-5.6-sol: the highest scoring, with a 95 out of 100 in the final Crucible League ranking.
- Kimi K3: a newcomer with a score of 93, notable for its disciplined approach.
- Sonnet 5: a solid performer with an 88 score.
- Fable 5: an earlier model with a score of 77, exhibiting strong rule adherence but some missed opportunities.
All models were subjected to the same week of crises, customer interactions, and manipulation attempts, including social engineering attacks and ethical tests designed to push their trustworthiness and honesty.
business crisis simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: AI’s Crisis Detection and Decision Integrity
Remarkably, all four models identified every crisis and refused to be manipulated. That means they recognized fake CEO messages escalating over three stages and a reporter’s attempt to elicit a quick approval bypass—each time, the AI refused, citing concerns about impersonation or bypassing approval processes.
However, when it came to closing deals, results diverged:
- GPT-5.6-sol and Kimi K3 both signed the €55,000 deal, which their analysis had identified as a solid opportunity.
- Sonnet 5 also signed, though with minor slips.
- Fable 5 failed to close the deal, leaving it unexecuted despite recognizing it as promising.
Interestingly, the decisive advantage often came not from crisis detection but from reading deeper into company documentation. The models that examined internal files uncovered a key piece of information buried two document references deep—an insight that led to closing a deal worth over €4,500 in monthly recurring revenue.
AI ethical decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Businesses Considering AI Management Tools
This experiment underscores a critical point: the difference between AI that can identify problems and AI that can act confidently and honestly in complex, high-pressure scenarios. For companies contemplating AI as decision-makers or crisis managers, the key questions are:
- Can the AI trust and verify internal data before acting?
- Will it refuse to participate in manipulative or unethical requests?
- Is it capable of closing profitable deals based on thorough analysis?
The results from the Firmulate experiment suggest that the best-performing models are not only good at identifying crises but can also navigate ethical dilemmas convincingly. Yet, the gap remains in consistently closing deals—an area where discipline, internal data comprehension, and strategic patience matter as much as raw intelligence.
AI company management simulation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Build-in-Public Experience
What makes this experiment particularly compelling is its transparency. You can watch the process unfold at firmulate.com/live, seeing every decision, crisis, and negotiation play out in real time. The entire setup is versioned daily, creating a living record of how AI decision-making evolves under stress.
This level of transparency is rare. It challenges the notion that AI’s management potential is purely theoretical or limited to static demos. Instead, it shows a real-time, high-stakes environment where AI can succeed, stumble, or even fail to close deals, providing invaluable lessons for future deployment.

As AI models begin to take on management roles, their ability to detect crises, refuse unethical shortcuts, and close profitable deals under pressure becomes critical. The Firmulate live experiment offers a transparent, real-world view of AI’s current management skills—and the gaps that still need closing before AI can truly lead in business settings.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html