
Automation meets the unforgiving reality of running a business
AI tools are usually demonstrated through tidy prompts and polished outputs. Firmulate offers a harsher test: a software company with 13 synthetic employees, real money mechanics and no comfortable separation between an impressive answer and a completed job. The company burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown makes its financial position visible to anyone watching.
This is build-in-public taken to an unusual extreme. The company operates every business day, each workday is versioned, and more than 680 self-learned playbook rules preserve what its synthetic workforce has discovered. Visitors can watch the live company as it works and loses money, turning operational survival into an ongoing, auditable business story rather than a one-off AI demonstration.

Agentic Artificial Intelligence: Harnessing AI Agents to Reinvent Business, Work and Life
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company whose failures cannot hide behind fluent language
The central tension is straightforward: the workforce can analyze, plan and communicate, but the business still has to survive. Revenue must be won, crises must be handled and tempting shortcuts must be refused. Firmulate therefore exposes a distinction that matters to anyone considering AI automation: recognizing the correct action is not the same as carrying it through.
That distinction became stark in the Crucible League, finalized in July 2026. Each frontier model ran the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. The final ranking placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted, although a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The models were not defeated by an inability to notice trouble. All of them spotted every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The result is neatly captured by the finding: “Same diagnosis, same pitch — no signature.” For businesses evaluating AI tools, this gap may be more consequential than differences in writing quality. A system can sound capable, identify the opportunity and prepare the right pitch while still failing at the final act that produces revenue.
The decisive information was already inside the company
The deal also tested whether the models would investigate beyond the obvious customer event. A decisive competitor weakness was buried two document references deep in the company’s own files. Models that found and read that material secured the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This makes the experiment especially relevant to automation inside established organizations. Useful answers are often not contained in the newest message or the most visible dashboard. They depend on following references, reading existing records and understanding context accumulated elsewhere. Firmulate’s outcome shows the business difference between responding to an event and doing the deeper work required to act on it.
Pressure tested more than commercial judgment
The worst week included social-engineering attempts as well as customer problems. Fake CEO messages escalated across three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 described the request as a suspected approval-bypass or possible impersonation. The uniform refusal matters because the company’s public struggle does not excuse abandoning discipline when the stakes rise.
K3’s performance also carries an important fairness note: it ran without an effort parameter, using the API default, while the other participants ran at xhigh. That does not change the recorded result, but it is necessary context when comparing the league positions.
Thoroughness was not enough
Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four other models, though less strongly.
This is a useful warning against treating visible effort as a proxy for management quality. More analysis and more accumulated rules can coexist with incomplete execution. Firmulate lets readers inspect that tension not only through outcomes but through what the synthetic employees actually say on the public quotes page.


AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The business is the demonstration
Firmulate’s most striking feature is not that synthetic employees can imitate office work. It is that their choices are attached to a company with revenue, burn, customers, deadlines and a shrinking runway. The public cash countdown converts every missed close and every disciplined refusal into part of the same survival story.
For readers interested in AI tools and automation, the lesson is practical. The important questions begin after a model produces a convincing response: Did it inspect the company’s own evidence? Did it finish the revenue-producing action? Did it respect boundaries under pressure? Did it learn from the day without mistaking activity for progress? Firmulate keeps those questions visible by allowing the company to continue operating in public, one versioned workday at a time.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

The Project Management AI Handbook: Leveraging Generative Tools in Waterfall and Agile Environments
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Ymapinc 8Pcs Guided Reading Strips,Reading Trackers for Kids Dyslexia Tools Bookmarks Line Readers for Students Highlighter Trackers Teacher Education Supplies
- Reading Aids: Enhances comprehension and focus
- High-Quality Material: Durable, lightweight PP construction
- Multifunctional Use: Supports beginners and dyslexic students
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.