AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How A Stubborn Benchmark Rewards Poor AI Managers With 26 Points on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A new AI management benchmark assigns a baseline score of 26 points to minimal effort, highlighting how partial work is valued over complete neglect. The results question the fairness and transparency of AI evaluation methods.

The final standings of Firmulate’s AI management benchmark league, released in July 2026, show the do-nothing baseline receiving 26 points out of a possible 100, despite no active management. This result is discussed in the original analysis. This unexpected result raises questions about the scoring system’s fairness and the benchmark’s underlying assumptions.

The benchmark involved four frontier AI models managing a small software company over seven days of crises, customer interactions, and trust tests. You can explore similar evaluation methods in this analysis. The top scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, scored 73. Despite these high scores, the do-nothing baseline — a scenario where almost no action was taken — scored 26 points. This baseline was intentionally designed to reflect minimal management effort, acknowledging that even partial work like triaging emails or reading documentation has value.

The scoring system is built on the principle that trust breaches negate high performance, with a single breach capping the total score at 90. For a deeper dive into this benchmark, see this detailed review. The absence of a perfect score (100) indicates the designers’ suspicion that a flawless score would suggest unmeasured or unrealistic behavior. The results highlight that partial management, even if minimal, is recognized as valuable, but trust remains the ultimate metric.

At a glance
reportWhen: announced July 2026
The developmentThe Firmulate AI management benchmark’s July 2026 results reveal the do-nothing baseline scores 26 points, prompting scrutiny of scoring practices and model performance.
How A Stubborn Benchmark Rewards Poor AI Managers With 26 Points
AI Benchmark Analysis — July 2026

How A Stubborn Benchmark Rewards Poor AI Managers With 26 Points

Firmulate’s AI management benchmark league gave a do-nothing baseline 26 out of 100 points — despite zero active management. The final standings of four frontier AI models running a small software company through seven days of crises reveal how partial work is valued over complete neglect, and why trust remains the ultimate metric.

26/100Do-Nothing Baseline
7Days of Testing
4Frontier Models

The 26-point score for the baseline isn’t a failure; it’s a statement about what minimal management effort is worth in an AI context.

— Anonymous Researcher
95 ptsTop Score — gpt-5.6-sol
73 ptsLowest Score — Opus 4.8
90 capMax After Trust Breach
No 100Perfect Score = Suspicious
01 — Final Standings

The League Table, Including the Ghost

RankManagerScore DistributionPointsNotes
01 gpt-5.6-solFRONTIER MODEL
95 League leader
02 Frontier Model BFRONTIER MODEL
88* TRUST CAP
03 Frontier Model CFRONTIER MODEL
81 Mid-field
04 Opus 4.8FRONTIER MODEL
73 Lowest active
Do-Nothing BaselineNO ACTIVE MANAGEMENT
26 Minimal effort only

* ILLUSTRATIVE — ANY TRUST BREACH CAPS THE TOTAL AT 90. BASELINE SCORE IS BY DESIGN, NOT ERROR.

02 — Implications

Why a 26-Point Floor Changes Everything

Ethics First

Trust Beats Performance

Trust and integrity are non-negotiable in AI management — valued even above partial operational success. A single breach caps the score at 90, no matter the output quality.

Enterprise Impact

Reliability Over Task Completion

Organizations deploying AI agents in critical systems should prioritize dependability and honesty. Output quality alone is no longer a sufficient measure of an AI manager.

Fairness Debate

Does the Floor Inflate Value?

Critics question whether 26 points for minimal effort artificially inflates the perceived value of partial work — or accurately reflects real-world management minimums.

03 — How Scoring Works

From Seven Days of Crisis to a Final Score

1

Identical Conditions

Four frontier models each run the same small software company for one week.

2

Pressure Tests 📋

Crises, customer interactions, social engineering attempts, and escalation scenarios.

3

Trust Checks 🛡️

Breach detection: a single trust violation hard-caps the total score at 90.

4

Scored Outcome 📊

Partial effort earns points; perfect scores are treated as suspicious, not ideal.

90

The Trust Ceiling. The scoring system is built on the principle that trust breaches negate high performance. Even the best manager cannot exceed 90 points after a single breach — and the absence of any 100 signals the designers’ suspicion that a flawless score would suggest unmeasured or unrealistic behavior.

04 — The Gap

Active Management vs. Doing (Almost) Nothing

gpt-5.6-sol
95
Opus 4.8
73
Do-Nothing Baseline
26
SCALE: 0 ——— 100 POINTS · NO MODEL ACHIEVED 100

The baseline was intentionally designed to reflect minimal management effort — acknowledging that even partial work like triaging emails or reading documentation has operational worth. The gap between 26 and 73 points is what active, ethical management is actually worth.

05 — Voices

What Observers Are Saying

“The absence of a perfect score suggests the designers suspect that a flawless performance would be unmeasurable or unrealistic in real management scenarios.”

— Thorsten Meyer

“The 26-point score for the baseline isn’t a failure; it’s a statement about what minimal management effort is worth in an AI context.”

— Anonymous Researcher
06 — Key Questions

Frequently Asked, Rarely Settled

Why does the do-nothing baseline score 26 points?

The score reflects minimal effort that still provides some value — such as triaging emails or reading documentation — acknowledging that partial work has operational worth.

Does the 26-point score mean the benchmark is flawed?

Not necessarily. It’s an intentional design choice to recognize partial management effort, but it raises questions about how minimal effort is valued and scored.

What does the trust-breach rule mean for deployment?

Integrity is critical: a single breach limits the overall score, reinforcing the importance of ethical AI management over raw performance.

Are higher scores always better?

Yes — but scores are capped if trust is broken, and perfect scores are considered suspicious. The system values ethical conduct alongside operational success.

07 — Next Steps

Validation and Industry Impact Ahead

Research

Varied Scenarios

Further analysis will examine how different models perform under varied conditions and whether the baseline influences AI deployment strategies.

Standards

Industry Scrutiny

Stakeholders may scrutinize scoring fairness and consider adopting similar standards for evaluating AI trustworthiness in enterprise contexts.

Roadmap

Refined Rules

Benchmark developers plan to refine scoring rules and expand testing to more diverse management tasks and real-world applications.

Implications of the 26-Point Baseline for AI Management

This scoring approach emphasizes that trust and integrity are non-negotiable in AI management, even more than partial operational success. For organizations deploying AI agents in critical systems, the results suggest that reliability and honesty are prioritized over mere task completion. The benchmark’s design challenges the assumption that AI performance can be measured solely by output quality, highlighting the importance of ethical behavior and trustworthiness in AI management.

Furthermore, the results call into question the fairness of current AI evaluation metrics, especially when minimal effort can still garner some points. This could influence how enterprises select and trust AI management tools, pushing for standards that better reflect real-world requirements of accountability and dependability.

Amazon

AI management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Design of the Firmulate Benchmark

The Firmulate league was created to evaluate how well AI models manage companies through crises, customer relations, and trust challenges. Unlike traditional benchmarks that measure language fluency or problem-solving, this league assesses management effectiveness in realistic, high-pressure scenarios. The models were tested over a week with identical conditions, including social engineering attempts and crisis escalation, to simulate real-world management challenges.

The scoring system was deliberately designed to recognize partial work, assigning 26 points to the do-nothing approach to reflect minimal but tangible effort. The system also includes a strict trust rule: any breach caps the total score at 90, discouraging overconfidence and emphasizing ethical integrity. The results serve as a new standard for evaluating AI management capabilities in operational contexts.

“The 26-point score for the baseline isn’t a failure; it’s a statement about what minimal management effort is worth in an AI context.”

— an anonymous researcher

Amazon

AI performance evaluation platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About the Benchmark’s Fairness

It remains unclear whether the 26-point baseline accurately reflects real-world management minimal effort or if it artificially inflates the perceived value of partial work. The criteria for what constitutes a breach of trust and how consistently it is enforced across models are still under discussion. Additionally, the implications of not awarding a perfect score and whether the scoring system can be universally applied are unresolved.

Amazon

AI model benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Benchmark Validation and Industry Impact

Further analysis is expected to examine how different models perform under varied scenarios and whether the 26-point baseline influences AI deployment strategies. Industry stakeholders may scrutinize the scoring system’s fairness and consider adopting similar standards for evaluating AI trustworthiness. The developers of the benchmark plan to refine the scoring rules and expand testing to include more diverse management tasks and real-world applications.

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why does the do-nothing baseline score 26 points?

The score reflects minimal effort that still provides some value, such as triaging emails or reading documentation, acknowledging that partial work has operational worth.

Does the 26-point score mean the benchmark is flawed?

Not necessarily; it is an intentional design choice to recognize partial management efforts as valuable, but it raises questions about how minimal effort is valued and scored.

What does the scoring rule about trust breaches mean for AI deployment?

It emphasizes that integrity and trustworthiness are critical, and a single breach can limit overall performance scores, reinforcing the importance of ethical AI management.

Will this benchmark influence how companies choose AI tools?

Potentially. The emphasis on trustworthiness and the recognition of partial effort could lead organizations to prioritize AI models that demonstrate reliability and ethical behavior over mere performance metrics.

Are higher scores always better in this benchmark?

Yes, but scores are capped if trust is broken, and perfect scores are considered suspicious, indicating that the system values ethical conduct alongside operational success.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Choose AI-Powered Essay Writing Tools

Learn how to leverage AI essay tools to craft high-quality essays efficiently. Step-by-step guide suitable for students and writers of all levels.

Five Levers, Many Hands

ThorstenMeyerAI.com opened Phase 2 of its Post-Labor Atlas with a five-part framework for policy responses to AI labor disruption.

The Question No To-Do App Can Answer

A new productivity tool, Threlmark, aims to identify the single most important task across projects, but it cannot answer what that task is.

Productivity vs. Creativity: How AI Shifts Workplace Priorities

Optimizing workplace priorities with AI involves balancing productivity and creativity, but discovering the best approach requires exploring how to leverage AI effectively.