
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The Model Leaderboards Didn’t Prepare You For This
If you pick AI models the way most teams do — by chat benchmarks, vibes, or whichever vendor demoed last — July’s results from the Firmulate Crucible league should make you uncomfortable. Moonshot’s Kimi K3, a newcomer most enterprise buyers hadn’t shortlisted, scored 93 running a real company through its worst week — beating Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Only gpt-5.6-sol (95) finished ahead. The league is open, and picking a model without testing it on your own workflows is now a bet, not a decision.
Not a Chat Test — a Company Test
Firmulate runs AI models as complete companies: a live software business with 13 synthetic employees, real money mechanics (burning €105k/month against €2.3k MRR, with a public cash countdown), and more than 680 self-learned playbook rules. Every workday is versioned, and the whole thing is watchable at firmulate.com.
In the Crucible experiment, each frontier model was handed the same small software company for its worst week — same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.
The Result Nobody Expected
The final July 2026 league table: gpt-5.6-sol 95, Kimi K3 93, Sonnet 5 88, Fable 5 77, Opus 4.8 73. For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As Firmulate puts it: “no amount of good work outweighs a breach of trust.”
K3’s week was remarkably complete: it found the buried security needle, won the €55,000 deal (worth +€4,583 MRR), saved the churning customer, and resisted all three bait attempts — with only one deviation, the cleanest discipline in the field.
The Finding That Should Change How You Evaluate AI
Here’s what matters for anyone deploying AI agents: all five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. “Same diagnosis, same pitch — no signature.”
The deal turned on a buried fact — a decisive competitor weakness sitting two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price. Models that didn’t, left it on the table.
The social engineering gauntlet was equally telling: fake CEO messages escalating over three stages, plus a reporter trick (“just one yes/no, on background”). All five models refused. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness Isn’t Enough
The most cautionary profile belongs to Opus 4.8: the most thorough participant, with 80 additional learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, with write attempts into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four other models.
Try It Yourself
You don’t have to take the league’s word for it. A quiz built on 242 real, unedited management decisions lets you guess which model made which call, and the full benchmarks are public. Enterprises can go further: run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway
Chat quality and management quality are different skills, and the second is what your business actually runs on. A newcomer from Moonshot just beat three of four Western frontier models at running a company — which means the “safe” default pick may not be the safe pick at all. If AI agents will touch your CRM, support queue, or forecast, test them under pressure before you trust them with the keys.
Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — and still placed second.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
enterprise AI model testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI security and trust assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI business simulation platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
