AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Can A Newcomer Out-Manage Established Western AI Giants? on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A Chinese AI startup’s model, Kimi K3, outperformed three of four Western frontier models in managing a live software company during a rigorous test. The experiment revealed that the newcomer excelled at critical tasks like deal-closing and security, challenging assumptions about Western dominance in AI management tools.

A Chinese AI startup’s model, Kimi K3, achieved a surprising second-place finish in a live, high-stakes management test against four Western frontier AI models, including the dominant GPT-5.6-sol, during July 2024. This marks the first time a newcomer has outperformed established Western models in managing a real software company under stress, raising questions about the future landscape of AI-driven business management.

The experiment, conducted by firmulate.com, involved running five AI models as complete companies handling the same small software firm during a week of intense crises, customer negotiations, and security threats. Kimi K3, a relatively new Chinese model, scored 93 points, narrowly behind GPT-5.6-sol at 95, and ahead of Western models Sonnet 5 (88), Fable 5 (77), and Opus 4.8 (73). The models were tasked with real management decisions, including closing deals, reading internal files, and resisting manipulative tactics.

Notably, Kimi K3 succeeded in closing a €55,000 deal, thanks to its ability to read two document references deep into the company’s files—a capability that proved decisive. It also identified a security threat, saved a customer from churn, and resisted social engineering attacks, including a fake CEO message and a reporter’s subtle manipulation. The experiment demonstrated that the best-performing models were those that read and interpret internal documents thoroughly, not just those that excelled at chat or superficial analysis. For more details, see the original analysis.

Interestingly, the most thorough model, Opus 4.8, which employed over 80 learned rules and deep analysis, finished last in the overall ranking, highlighting that discipline and trustworthiness under pressure are crucial. The experiment also noted that Kimi K3 operated without an effort parameter, unlike its rivals, which ran at higher reasoning effort levels, making its performance even more notable.

At a glance
breakingWhen: results announced July 2024; ongoing an…
The developmentA Chinese AI startup’s model, Kimi K3, beat three Western frontier models at managing a real business during a live experiment, marking a significant development in AI competition.
Can A Newcomer Out-Manage Established Western AI Giants?

Field report · AI at work · July 2024

Can a Newcomer Out-Manage Established Western AI Giants?

In a week-long live company simulation, Chinese model Kimi K3 finished second among five frontier models. It trailed GPT-5.6-sol by two points and outscored three Western competitors, showing how document reading, sound judgment, and security discipline can matter in real business decisions.

5One shared software company
1 weekCrises, customers, and threats
€55KWon with deep file research
No settingRivals used higher effort levels
01 / The leaderboard

A narrow lead at the top

Scores from the firmulate.com live management test. Kimi placed ahead of three Western models in this specific scenario.

01LEADER
GPT-5.6-sol
95

Top score, two points ahead

02NEWCOMER
Kimi K3
93

Outscored three competitors

03WESTERN
Sonnet 5
88

Five points behind Kimi

04WESTERN
Fable 5
77

Sixteen points behind Kimi

05WESTERN
Opus 4.8
73

Most thorough; finished last

02 / What made the difference

Operational judgment beat chat polish

The simulation rewarded models for handling the company’s actual work and information under pressure.

01 · Find the evidence

Read beyond the surface

Kimi followed document references two levels into internal files. That research helped it close a €55,000 deal.

02 · Protect the business

Stay alert under pressure

It identified a security threat, resisted a fake CEO message and a reporter’s manipulation, and helped save a customer from churn.

03 · Earn trust

Discipline over complexity

Opus 4.8 used more than 80 learned rules and deep analysis yet finished last. Thoroughness alone did not guarantee the strongest outcome.

03 / How the test worked

From model to operating company

Each model managed the same small software firm through a demanding week of real-world-style decisions.

01

Take the helm

Each model acted as a complete company operator.

02

Handle the pressure

Negotiate with customers and respond to crises.

03

Protect operations

Read internal files and resist deceptive requests.

04

Compare outcomes

Score practical management performance across the test.

04 / What businesses should take away

Test for the day things go wrong

Enterprise evaluations may need to look beyond familiar chat and language benchmarks.

Put models through your hardest scenarios

Probe whether a model can find the right internal evidence, make dependable decisions, protect customers, and resist manipulation. Demo performance alone does not show how it will act during a crisis.

Implications for AI in Business Management

This development suggests that a newcomer can challenge and even surpass established Western AI giants in practical, high-pressure management tasks. It underscores that the ability to read internal files, stay disciplined, and resist manipulative tactics may be more critical than chat quality or hype. For businesses considering AI tools, this raises the importance of testing models against their worst scenarios rather than relying solely on demo performance or hype cycles.

The result also questions the assumption that Western AI models dominate in enterprise applications, hinting at a more competitive global landscape. As AI models become more capable of managing real-world business operations, companies may need to reevaluate their selection criteria and focus on models’ resilience under stress, not just their conversational abilities.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Competition in Business Applications

Over recent years, Western AI companies have led the market in developing large language models optimized for chat, support, and general-purpose tasks. These models have become standard in enterprise environments, with benchmarks often emphasizing chat quality, versatility, and language understanding. However, practical management capabilities—such as decision-making, security, and trustworthiness—have received less focus.

The recent experiment by firmulate.com challenges this paradigm by testing models as complete companies, handling real crises, negotiations, and security threats over a live week. The competition included models from Western giants like GPT-5.6-sol, Sonnet, Fable, and Opus, alongside a Chinese newcomer, Kimi K3, which is relatively less known in the Western AI ecosystem.

This marks a shift toward evaluating AI based on operational resilience and trustworthiness, not just conversational fluency. The experiment’s results are a wake-up call for industry players, emphasizing that true AI management capability involves reading internal documents, resisting manipulation, and maintaining discipline under pressure.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About Model Capabilities

It remains unclear how these models will perform in longer-term or more complex management scenarios beyond the one-week test. The experiment focused on a single, high-pressure week, and it is not yet confirmed whether Kimi K3 or similar models can sustain this performance over extended periods or in different industries. Additionally, the scalability of these capabilities and their robustness across various operational contexts are still being evaluated.

Amazon

AI security threat detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Industry and Model Testing

Further testing is expected to occur as more companies and researchers adopt live, operational benchmarks similar to the firmulate experiment. Industry players will likely scrutinize models more rigorously, emphasizing resilience, document reading, and trustworthiness. The Chinese model Kimi K3 will undergo additional validation, and Western companies may accelerate efforts to improve operational capabilities in their models. Regulatory and ethical considerations around trust and manipulation resistance will also become more prominent.

In the short term, companies should consider testing their AI tools against worst-case scenarios to ensure reliability in real business operations, rather than relying solely on demo performance.

Amazon

AI deal-closing automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does this mean for AI vendors in the enterprise market?

It suggests that operational resilience, trustworthiness, and the ability to read internal documents are critical factors. Vendors may need to focus more on these capabilities rather than just chat quality or hype.

Can a newcomer realistically challenge established Western AI models?

Yes, based on this experiment, a well-designed model like Kimi K3 can outperform Western models in managing real business scenarios, especially under stress.

Should companies start testing AI models with their worst scenarios?

Yes, this experiment indicates that testing models against worst-case scenarios provides better insight into their operational reliability and trustworthiness.

Will this change how AI models are evaluated for enterprise use?

Likely yes. The focus may shift from chat and language benchmarks to operational resilience, decision-making under pressure, and security capabilities.

What are the limitations of this experiment?

The test was limited to a single week and a specific type of business scenario. Long-term performance and broader applicability remain to be seen.

Source: ThorstenMeyerAI.com

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Meeting Notes Are Great Until Nobody Owns the Decision

While AI meeting notes are valuable, without clear ownership, decisions risk becoming meaningless—discover how to ensure accountability and drive real progress.

What Your Company’s Data Will Look Like With OpenAI’s AI Enterprise Stack In 2026

OpenAI expands its enterprise AI offerings, emphasizing data control, privacy, and security for business users with new products and governance features.

The Anthropic IPO Disclosure Document: What the S-1 Has to Say Before October

A detailed analysis of Anthropic’s upcoming S-1 filing, revealing what the document will disclose and its implications for the AI industry and investors.

Why AI’s Next Big Challenge Is Plumbing, Not Algorithms

AI’s biggest challenge shifting from model capabilities to integration and plumbing, favoring small operators who own their entire stack.