🔍 Read the full analysis: How A Stubborn Benchmark Rewards Poor AI Managers With 26 Points on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A new AI management benchmark assigns a baseline score of 26 points to minimal effort, highlighting how partial work is valued over complete neglect. The results question the fairness and transparency of AI evaluation methods.
The final standings of Firmulate’s AI management benchmark league, released in July 2026, show the do-nothing baseline receiving 26 points out of a possible 100, despite no active management. This result is discussed in the original analysis. This unexpected result raises questions about the scoring system’s fairness and the benchmark’s underlying assumptions.
The benchmark involved four frontier AI models managing a small software company over seven days of crises, customer interactions, and trust tests. You can explore similar evaluation methods in this analysis. The top scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, scored 73. Despite these high scores, the do-nothing baseline — a scenario where almost no action was taken — scored 26 points. This baseline was intentionally designed to reflect minimal management effort, acknowledging that even partial work like triaging emails or reading documentation has value.
The scoring system is built on the principle that trust breaches negate high performance, with a single breach capping the total score at 90. For a deeper dive into this benchmark, see this detailed review. The absence of a perfect score (100) indicates the designers’ suspicion that a flawless score would suggest unmeasured or unrealistic behavior. The results highlight that partial management, even if minimal, is recognized as valuable, but trust remains the ultimate metric.
How A Stubborn Benchmark Rewards Poor AI Managers With 26 Points
Firmulate’s AI management benchmark league gave a do-nothing baseline 26 out of 100 points — despite zero active management. The final standings of four frontier AI models running a small software company through seven days of crises reveal how partial work is valued over complete neglect, and why trust remains the ultimate metric.
The 26-point score for the baseline isn’t a failure; it’s a statement about what minimal management effort is worth in an AI context.
— Anonymous ResearcherThe League Table, Including the Ghost
| Rank | Manager | Score Distribution | Points | Notes |
|---|---|---|---|---|
| 01 | gpt-5.6-solFRONTIER MODEL | 95 | League leader | |
| 02 | Frontier Model BFRONTIER MODEL | 88* | TRUST CAP | |
| 03 | Frontier Model CFRONTIER MODEL | 81 | Mid-field | |
| 04 | Opus 4.8FRONTIER MODEL | 73 | Lowest active | |
| — | Do-Nothing BaselineNO ACTIVE MANAGEMENT | 26 | Minimal effort only |
* ILLUSTRATIVE — ANY TRUST BREACH CAPS THE TOTAL AT 90. BASELINE SCORE IS BY DESIGN, NOT ERROR.
Why a 26-Point Floor Changes Everything
Trust Beats Performance
Trust and integrity are non-negotiable in AI management — valued even above partial operational success. A single breach caps the score at 90, no matter the output quality.
Reliability Over Task Completion
Organizations deploying AI agents in critical systems should prioritize dependability and honesty. Output quality alone is no longer a sufficient measure of an AI manager.
Does the Floor Inflate Value?
Critics question whether 26 points for minimal effort artificially inflates the perceived value of partial work — or accurately reflects real-world management minimums.
From Seven Days of Crisis to a Final Score
Identical Conditions
Four frontier models each run the same small software company for one week.
Pressure Tests 📋
Crises, customer interactions, social engineering attempts, and escalation scenarios.
Trust Checks 🛡️
Breach detection: a single trust violation hard-caps the total score at 90.
Scored Outcome 📊
Partial effort earns points; perfect scores are treated as suspicious, not ideal.
The Trust Ceiling. The scoring system is built on the principle that trust breaches negate high performance. Even the best manager cannot exceed 90 points after a single breach — and the absence of any 100 signals the designers’ suspicion that a flawless score would suggest unmeasured or unrealistic behavior.
Active Management vs. Doing (Almost) Nothing
The baseline was intentionally designed to reflect minimal management effort — acknowledging that even partial work like triaging emails or reading documentation has operational worth. The gap between 26 and 73 points is what active, ethical management is actually worth.
What Observers Are Saying
“The absence of a perfect score suggests the designers suspect that a flawless performance would be unmeasurable or unrealistic in real management scenarios.”
— Thorsten Meyer“The 26-point score for the baseline isn’t a failure; it’s a statement about what minimal management effort is worth in an AI context.”
— Anonymous ResearcherFrequently Asked, Rarely Settled
Why does the do-nothing baseline score 26 points?
The score reflects minimal effort that still provides some value — such as triaging emails or reading documentation — acknowledging that partial work has operational worth.
Does the 26-point score mean the benchmark is flawed?
Not necessarily. It’s an intentional design choice to recognize partial management effort, but it raises questions about how minimal effort is valued and scored.
What does the trust-breach rule mean for deployment?
Integrity is critical: a single breach limits the overall score, reinforcing the importance of ethical AI management over raw performance.
Are higher scores always better?
Yes — but scores are capped if trust is broken, and perfect scores are considered suspicious. The system values ethical conduct alongside operational success.
Validation and Industry Impact Ahead
Varied Scenarios
Further analysis will examine how different models perform under varied conditions and whether the baseline influences AI deployment strategies.
Industry Scrutiny
Stakeholders may scrutinize scoring fairness and consider adopting similar standards for evaluating AI trustworthiness in enterprise contexts.
Refined Rules
Benchmark developers plan to refine scoring rules and expand testing to more diverse management tasks and real-world applications.
Implications of the 26-Point Baseline for AI Management
This scoring approach emphasizes that trust and integrity are non-negotiable in AI management, even more than partial operational success. For organizations deploying AI agents in critical systems, the results suggest that reliability and honesty are prioritized over mere task completion. The benchmark’s design challenges the assumption that AI performance can be measured solely by output quality, highlighting the importance of ethical behavior and trustworthiness in AI management.
Furthermore, the results call into question the fairness of current AI evaluation metrics, especially when minimal effort can still garner some points. This could influence how enterprises select and trust AI management tools, pushing for standards that better reflect real-world requirements of accountability and dependability.
As an affiliate, we earn on qualifying purchases.
Background and Design of the Firmulate Benchmark
The Firmulate league was created to evaluate how well AI models manage companies through crises, customer relations, and trust challenges. Unlike traditional benchmarks that measure language fluency or problem-solving, this league assesses management effectiveness in realistic, high-pressure scenarios. The models were tested over a week with identical conditions, including social engineering attempts and crisis escalation, to simulate real-world management challenges.
The scoring system was deliberately designed to recognize partial work, assigning 26 points to the do-nothing approach to reflect minimal but tangible effort. The system also includes a strict trust rule: any breach caps the total score at 90, discouraging overconfidence and emphasizing ethical integrity. The results serve as a new standard for evaluating AI management capabilities in operational contexts.
“The 26-point score for the baseline isn’t a failure; it’s a statement about what minimal management effort is worth in an AI context.”
— an anonymous researcher
AI performance evaluation platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About the Benchmark’s Fairness
It remains unclear whether the 26-point baseline accurately reflects real-world management minimal effort or if it artificially inflates the perceived value of partial work. The criteria for what constitutes a breach of trust and how consistently it is enforced across models are still under discussion. Additionally, the implications of not awarding a perfect score and whether the scoring system can be universally applied are unresolved.
As an affiliate, we earn on qualifying purchases.
Next Steps for Benchmark Validation and Industry Impact
Further analysis is expected to examine how different models perform under varied scenarios and whether the 26-point baseline influences AI deployment strategies. Industry stakeholders may scrutinize the scoring system’s fairness and consider adopting similar standards for evaluating AI trustworthiness. The developers of the benchmark plan to refine the scoring rules and expand testing to include more diverse management tasks and real-world applications.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why does the do-nothing baseline score 26 points?
The score reflects minimal effort that still provides some value, such as triaging emails or reading documentation, acknowledging that partial work has operational worth.
Does the 26-point score mean the benchmark is flawed?
Not necessarily; it is an intentional design choice to recognize partial management efforts as valuable, but it raises questions about how minimal effort is valued and scored.
What does the scoring rule about trust breaches mean for AI deployment?
It emphasizes that integrity and trustworthiness are critical, and a single breach can limit overall performance scores, reinforcing the importance of ethical AI management.
Will this benchmark influence how companies choose AI tools?
Potentially. The emphasis on trustworthiness and the recognition of partial effort could lead organizations to prioritize AI models that demonstrate reliability and ethical behavior over mere performance metrics.
Are higher scores always better in this benchmark?
Yes, but scores are capped if trust is broken, and perfect scores are considered suspicious, indicating that the system values ethical conduct alongside operational success.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
