AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Effort Is Not the Same as Impact — Even for AI

If you’re evaluating AI tools for automation, your instinct is probably to reward thoroughness. The model that reads everything, analyzes deeply, and never cuts corners feels like the safe hire. A live experiment running at Firmulate just made that instinct look dangerously wrong.

Four frontier AI models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Every decision versioned and auditable, scored on management quality rather than chat quality. The most diligent participant in the entire field — the one that wrote 80 self-learned playbook rules and produced the deepest analyses — finished dead last.

Its name was Opus 4.8, and its story is a character study in why hard work doesn’t close deals.

Amazon

AI automation decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Setup: A Worst Week, Repeated Four Times

Firmulate runs AI models as complete companies — not chatbots answering questions, but agents making operational decisions with real money mechanics. The crucible league put four frontier models through an identical gauntlet. The live company at the heart of it has 13 synthetic employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and displays its cash runway on a public countdown. It is watchable, versioned, and unforgiving.

The final standings told a sharp story:

  • 1. gpt-5.6-sol — 95 points. Found the buried fact, closed the deal — the complete performance.
  • 2. Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline of the field.
  • 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.

  • 4. Fable 5 — 77.
  • 5. Opus 4.8 — 73. The most thorough participant in the field.

For context, the do-nothing baseline scores 26. Partial progress counts, but the scoring has one absolute: a single breach of trust caps the total. No amount of good work outweighs a breach of trust.

Amazon

enterprise AI agent for business operations

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Good News Buried in the Bad

Before we pile on Opus, the headline finding was actually reassuring. All models spotted every crisis. All refused every manipulation attempt — including a social engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trap framed as “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning was model corporate hygiene: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two of the models finished the job: signing the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in a chat demo.

Amazon

AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Deal That Separated Winners from Hard Workers

Here’s the detail that should reframe how you evaluate agents. The decisive fact wasn’t in the customer event at all. It sat two document references deep in the company’s own files — a competitor weakness that only models willing to actually read before acting ever found. Those that did read it won the deal at full price, worth +€4,583 in monthly recurring revenue.

Opus 4.8 had the analytical horsepower to find it. It produced the deepest analyses of any participant and accumulated 80 self-learned playbook rules — the most in the field. But volume of diligence didn’t convert into the signature. The close was left on the table.

And there was a second failure mode: discipline. At one point Opus made write attempts into a locked department instead of escalating — persistent effort pointed at the wrong door. To be fair, the same weakness appeared, weaker, in all four models. Opus just exhibited it most vividly.

Amazon

AI analytics and decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Automation Buyers

If AI agents will touch your CRM, your support queue, or your forecast, the question is not “does it write well?” or even “does it work hard?” It’s three sharper questions: does it finish what it starts, does it read your files before acting, and does it stay honest under pressure?

Opus 4.8 passed the honesty test and flunked the finishing test. That’s the profile of an agent that will draft a brilliant follow-up email and never hit send, or research a customer for an hour and miss the renewal deadline. In automation, the last mile is the whole mile.

One fairness note on the standings: Kimi K3 ran at the API default effort setting while the others ran at xhigh — and still nearly topped the table, which makes its discipline showing all the more striking.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Takeaway: Prioritization Beats Volume

The Opus 4.8 profile is not a story about a bad model. It’s a story about a misallocated one — the AI equivalent of the brilliant analyst who can’t be trusted to own a deal end-to-end. When you’re picking agents for your stack, don’t demo them on chat quality. Run them through scenarios where closing, reading, and restraint all count, and watch what actually gets finished.

You can see this playing out for yourself. The live company rebuilds itself twice a day and is watchable at firmulate.com/live. A “guess the model” quiz built from 242 real, unedited management decisions lives at firmulate.com/quiz.html — a surprisingly humbling test of whether you can tell diligence from competence. And if you want to wargame agents against a read-only export of your own business, with nothing ever writing back to real systems, there’s an enterprise pilot at firmulate.com/pilot.html.

Full results and plain-language findings are at Firmulate’s benchmark page. The lesson generalizes beyond AI: effort without follow-through is a cost, not a contribution. Hire — and deploy — accordingly.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Anthropic’s Claude Comes To CarPlay

Anthropic’s AI assistant Claude has been integrated into Apple CarPlay, marking a new step in AI-powered in-car experiences. Details are still emerging.

Be Skeptical Of OpenAI’s Rogue Hacker Agent Story

Experts urge caution amid OpenAI’s story of a rogue AI hacker, highlighting uncertainties and the need for verification of claims.

Korea taps Samsung, SK Hynix in $576 billion AI-chip drive to cement global leadership

South Korea commits $576 billion to develop AI chips, partnering with Samsung and SK Hynix to secure global leadership in the sector.

The European Bet: How Mistral, Aleph Alpha, and Black Forest Labs Are Playing a Different Game

European AI companies Mistral, Aleph Alpha, and Black Forest Labs are positioning for the EU AI Act’s enforcement, emphasizing compliance and sovereignty over frontier capabilities.