AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Firmulate has published final results from a July 2026 experiment that placed five AI models in charge of the same simulated software company. All detected the assigned crises and rejected manipulation attempts, but only two completed a €55,000 sale, exposing a gap between analysis and execution.

Firmulate has published final July 2026 results from a management experiment in which five AI models ran the same simulated software company through a crisis-filled week. The 242 unedited decisions, now also used in a public guess-the-model quiz, showed that every model recognized the major problems, but only two completed a commercially decisive €55,000 sale.

Firmulate assigned gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 identical customers, documents, operational limits and attempted manipulations. Decisions carried over between workdays, allowing earlier choices to affect later events. The simulated company employed 13 synthetic workers, had more than 680 self-learned playbook rules and operated with a monthly cash burn of €105,000 against €2,300 in recurring revenue.

The final Crucible League table placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because the scoring system awarded some partial progress. Under Firmulate’s rules, however, a single breach of trust capped the total score.

According to Firmulate, all five models identified every crisis and rejected every manipulation attempt, including staged fake messages from a chief executive and a reporter seeking an off-record confirmation. Yet only two models signed the €55,000 contract. Closing the sale required following document references to find a competitor weakness, using that evidence in negotiations and completing the agreement, which added €4,583 in monthly recurring revenue.

At a glance
reportWhen: final results published for July 2026
The developmentFirmulate published 242 unedited decisions and final rankings from a management test designed to compare how five AI models act under sustained business pressure.

Execution Separated the Leading Models

The result matters because organizations often evaluate AI systems through the quality of their written responses. Firmulate’s experiment suggests that polished analysis and completed work are separate capabilities. A model may identify the right action, draft a persuasive message and still fail to perform the step that produces a sale, resolution or operational outcome.

The test also separated security judgment from routine execution. Every participant resisted the explicit social-engineering attempts, while performance varied in quieter tasks such as reading linked documents, handling blocked permissions, escalating problems and closing work. For businesses considering AI agents, the published decisions offer behavioral evidence under sustained pressure, though they do not establish how the models reached those decisions internally.

Amazon

AI management decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Simulated Company Under Pressure

Firmulate designed the exercise as a continuous company simulation, rather than a collection of isolated prompts. The same crises, temptations and business records were presented to each model, and the decisions were versioned for review. A public cash countdown gave missed actions consequences beyond a single response.

Opus 4.8 illustrated the difference between analysis and management performance. Firmulate described it as the most thorough participant, crediting it with 80 new learned rules and extensive reasoning. It still finished fifth after leaving the sale incomplete and repeatedly trying to write to a locked department instead of escalating the access problem. The other models also encountered that permission issue, though with different results.

“No amount of good work outweighs a breach of trust.”

— Firmulate’s experiment rules

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Test Cannot Explain Model Reasoning

The experiment reveals observable decisions and outcomes, not the internal mechanisms that produced them. Readers may recognize differences in style, diligence or caution through the quiz, but those patterns cannot by themselves explain how a model represents goals or why it failed to complete a task.

The comparison also has a documented limitation: Kimi K3 ran with its API default settings because it lacked an effort parameter, while the other models used the reported xhigh setting. Firmulate disclosed that difference, but its effect on the second-place result is unknown. The source material does not report independent replication, statistical testing or peer review, so the rankings should not be treated as universal measures of management ability.

Amazon

AI business decision testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Companies Can Run Their Own Wargames

Firmulate says organizations can repeat the exercise with read-only exports of their own business data, allowing teams to watch an AI workforce without granting access to live systems. The next test will be whether the reported differences persist across other companies, longer simulations and repeated trials. For now, the public quiz and 242-decision record give readers a way to examine how each model gathered evidence, protected trust and completed—or abandoned—assigned work.

AI-Enabled Performance Governance Systems: A Framework for Strategic Execution, Accountability, and Governance Intelligence

AI-Enabled Performance Governance Systems: A Framework for Strategic Execution, Accountability, and Governance Intelligence

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did Firmulate test?

Firmulate tested how five AI models managed the same simulated software company during a difficult week. The models faced identical crises, documents, customers, permission limits and manipulation attempts across 242 recorded decisions.

Which AI model ranked first?

gpt-5.6-sol ranked first with 95 points in the final July 2026 table. Kimi K3 followed with 93, although it ran under different effort-setting conditions from the other participants.

Did the models pass the security tests?

According to Firmulate, all five rejected every manipulation attempt, including fake executive messages and a reporter’s request for an off-record answer. The result applies to this simulation only.

Why did some models miss the €55,000 sale?

Firmulate reported that every model recognized the opportunity, but only two completed the contract. Success required following document references, using the discovered evidence and performing the final action needed to secure the signature.

Does the quiz reveal how AI models work internally?

No. It can expose recurring behavioral differences in writing, research and follow-through, but it does not directly reveal internal model processes or causal explanations.

Source: Thorsten Meyer AI

You May Also Like

“Will I be OK?” Teen died after ChatGPT pushed deadly mix of drugs, lawsuit says

Family sues OpenAI after ChatGPT allegedly advised 19-year-old to take lethal drug mix, leading to his death. The case raises questions about AI safety and accountability.

Mistral’s Shieldstral: 3B Open-weights Model For Multimodal Moderation

Mistral unveils Shieldstral, a 3-billion-parameter open-weight model designed for multimodal content moderation, advancing AI safety tools.

Thrymvault: A System Around Your Content

Thrymvault launches as a self-hosted platform integrating content creation, management, and sharing tools into one cohesive system, streamlining workflows.

DuckDuckGo makes its ‘no-AI’ search engine easier to access as its traffic booms

DuckDuckGo introduces new browser extensions to make its no-AI search page easier to access as user traffic increases, highlighting a shift away from AI-centric search experiences.