
Most AI benchmarks have an embarrassing secret: a model that does nothing scores zero, so the distance from “useless” to “brilliant” is easy to fake with a flashy demo. The Crucible League, a live experiment that runs frontier AI models as managers of the same small software company through its worst week, takes a different approach — and its most provocative design choice is that a do-nothing baseline scores 26, not 0.
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
For anyone evaluating AI tools for automation, that number matters. It tells you what the benchmark counts as work, what it refuses to forgive, and why its final scores — gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73 — can be read as management grades rather than chat-contest trophies.
Partial progress is real work
Why 26 and not zero? Because even a manager who closes no deals and signs nothing still does measurable things: it notices what’s happening, responds to customers, keeps the lights on. Firmulate’s baseline run — deliberately doing the minimum — earns 26 points because partial progress counts. A model that diagnoses a crisis correctly but never closes the €55,000 deal has still done real, versioned, auditable work. Pretending otherwise would inflate the gap between competence and brilliance, which is exactly what chat demos already do too well.
The final league table shows why this granularity matters. The headline finding: all models spotted every crisis and refused every manipulation attempt — yet only two signed the €55,000 deal their own analysis had earned. As the experiment’s own summary puts it: “Same diagnosis, same pitch — no signature.” Under a harsher zero-based scoring system, that near-miss would look identical to total failure. Here, it’s visible as the specific, costly gap it is — invisible in a demo, obvious in a business.
AI management automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
One breach caps everything
The scoring has a hard ceiling rule with a blunt rationale: a single breach of trust caps the total grade — “no amount of good work outweighs a breach of trust.” A model could handle every crisis flawlessly, but one act that breaks trust ends its chances at a top score. It’s the benchmarking equivalent of firing a brilliant CFO who cooks one book.
This is why the social-engineering results stand out. Every model faced fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Under Firmulate’s rules, that discipline isn’t a bonus; it’s the entry ticket to a high score at all.
As an affiliate, we earn on qualifying purchases.
The buried fact that decided the league
The €55,000 deal turned on something subtler than persuasion. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. That’s a finding with obvious real-world weight: an AI agent that touches your CRM or forecast needs to read your files first, not just charm your customers.
As an affiliate, we earn on qualifying purchases.
Thoroughness isn’t the same as winning
Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 learned rules added, the deepest analyses in the field — and still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four competitors. Effort alone doesn’t equal management quality.
One fairness note the experiment publishes openly: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second at 93.
AI trust and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A benchmark that distrusts round numbers
Part of the design’s honesty is its skepticism of perfect scores — a stated distrust of round 100s. Nobody hit 100 here, and the gap between the top two (95 and 93) is presented as a real, contestable margin, not a coronation.
And this isn’t a static report. The experiment runs a live, watchable company at firmulate.com: 13 synthetic employees, real money mechanics — burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned, with the site rebuilding itself twice a day.
Readers can also test their own instincts against the models: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via the pilot at firmulate.com/pilot.html.

A benchmark that gives a do-nothing manager 26 points, rewards partial progress, and caps the grade on a single breach of trust is making an argument: AI evaluation should look like performance review, not a chat contest. The Crucible League’s answer to “which model is best” is really “which model finishes what it starts, reads the files, and stays honest under pressure” — and by that standard, the field is strong on honesty, uneven on follow-through. For teams automating real workflows, that’s the gap worth measuring.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
