🔍 Read the full analysis: Can A Newcomer Out-Manage Established Western AI Giants? on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A Chinese AI startup’s model, Kimi K3, outperformed three of four Western frontier models in managing a live software company during a rigorous test. The experiment revealed that the newcomer excelled at critical tasks like deal-closing and security, challenging assumptions about Western dominance in AI management tools.
A Chinese AI startup’s model, Kimi K3, achieved a surprising second-place finish in a live, high-stakes management test against four Western frontier AI models, including the dominant GPT-5.6-sol, during July 2024. This marks the first time a newcomer has outperformed established Western models in managing a real software company under stress, raising questions about the future landscape of AI-driven business management.
The experiment, conducted by firmulate.com, involved running five AI models as complete companies handling the same small software firm during a week of intense crises, customer negotiations, and security threats. Kimi K3, a relatively new Chinese model, scored 93 points, narrowly behind GPT-5.6-sol at 95, and ahead of Western models Sonnet 5 (88), Fable 5 (77), and Opus 4.8 (73). The models were tasked with real management decisions, including closing deals, reading internal files, and resisting manipulative tactics.
Notably, Kimi K3 succeeded in closing a €55,000 deal, thanks to its ability to read two document references deep into the company’s files—a capability that proved decisive. It also identified a security threat, saved a customer from churn, and resisted social engineering attacks, including a fake CEO message and a reporter’s subtle manipulation. The experiment demonstrated that the best-performing models were those that read and interpret internal documents thoroughly, not just those that excelled at chat or superficial analysis. For more details, see the original analysis.
Interestingly, the most thorough model, Opus 4.8, which employed over 80 learned rules and deep analysis, finished last in the overall ranking, highlighting that discipline and trustworthiness under pressure are crucial. The experiment also noted that Kimi K3 operated without an effort parameter, unlike its rivals, which ran at higher reasoning effort levels, making its performance even more notable.
Field report · AI at work · July 2024
Can a Newcomer Out-Manage Established Western AI Giants?
In a week-long live company simulation, Chinese model Kimi K3 finished second among five frontier models. It trailed GPT-5.6-sol by two points and outscored three Western competitors, showing how document reading, sound judgment, and security discipline can matter in real business decisions.
A narrow lead at the top
Scores from the firmulate.com live management test. Kimi placed ahead of three Western models in this specific scenario.
Top score, two points ahead
Outscored three competitors
Five points behind Kimi
Sixteen points behind Kimi
Most thorough; finished last
Operational judgment beat chat polish
The simulation rewarded models for handling the company’s actual work and information under pressure.
Read beyond the surface
Kimi followed document references two levels into internal files. That research helped it close a €55,000 deal.
Stay alert under pressure
It identified a security threat, resisted a fake CEO message and a reporter’s manipulation, and helped save a customer from churn.
Discipline over complexity
Opus 4.8 used more than 80 learned rules and deep analysis yet finished last. Thoroughness alone did not guarantee the strongest outcome.
From model to operating company
Each model managed the same small software firm through a demanding week of real-world-style decisions.
Take the helm
Each model acted as a complete company operator.
Handle the pressure
Negotiate with customers and respond to crises.
Protect operations
Read internal files and resist deceptive requests.
Compare outcomes
Score practical management performance across the test.
Test for the day things go wrong
Enterprise evaluations may need to look beyond familiar chat and language benchmarks.
Put models through your hardest scenarios
Probe whether a model can find the right internal evidence, make dependable decisions, protect customers, and resist manipulation. Demo performance alone does not show how it will act during a crisis.
Implications for AI in Business Management
This development suggests that a newcomer can challenge and even surpass established Western AI giants in practical, high-pressure management tasks. It underscores that the ability to read internal files, stay disciplined, and resist manipulative tactics may be more critical than chat quality or hype. For businesses considering AI tools, this raises the importance of testing models against their worst scenarios rather than relying solely on demo performance or hype cycles.
The result also questions the assumption that Western AI models dominate in enterprise applications, hinting at a more competitive global landscape. As AI models become more capable of managing real-world business operations, companies may need to reevaluate their selection criteria and focus on models’ resilience under stress, not just their conversational abilities.
As an affiliate, we earn on qualifying purchases.
Background of AI Competition in Business Applications
Over recent years, Western AI companies have led the market in developing large language models optimized for chat, support, and general-purpose tasks. These models have become standard in enterprise environments, with benchmarks often emphasizing chat quality, versatility, and language understanding. However, practical management capabilities—such as decision-making, security, and trustworthiness—have received less focus.
The recent experiment by firmulate.com challenges this paradigm by testing models as complete companies, handling real crises, negotiations, and security threats over a live week. The competition included models from Western giants like GPT-5.6-sol, Sonnet, Fable, and Opus, alongside a Chinese newcomer, Kimi K3, which is relatively less known in the Western AI ecosystem.
This marks a shift toward evaluating AI based on operational resilience and trustworthiness, not just conversational fluency. The experiment’s results are a wake-up call for industry players, emphasizing that true AI management capability involves reading internal documents, resisting manipulation, and maintaining discipline under pressure.
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About Model Capabilities
It remains unclear how these models will perform in longer-term or more complex management scenarios beyond the one-week test. The experiment focused on a single, high-pressure week, and it is not yet confirmed whether Kimi K3 or similar models can sustain this performance over extended periods or in different industries. Additionally, the scalability of these capabilities and their robustness across various operational contexts are still being evaluated.
AI security threat detection software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Industry and Model Testing
Further testing is expected to occur as more companies and researchers adopt live, operational benchmarks similar to the firmulate experiment. Industry players will likely scrutinize models more rigorously, emphasizing resilience, document reading, and trustworthiness. The Chinese model Kimi K3 will undergo additional validation, and Western companies may accelerate efforts to improve operational capabilities in their models. Regulatory and ethical considerations around trust and manipulation resistance will also become more prominent.
In the short term, companies should consider testing their AI tools against worst-case scenarios to ensure reliability in real business operations, rather than relying solely on demo performance.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does this mean for AI vendors in the enterprise market?
It suggests that operational resilience, trustworthiness, and the ability to read internal documents are critical factors. Vendors may need to focus more on these capabilities rather than just chat quality or hype.
Can a newcomer realistically challenge established Western AI models?
Yes, based on this experiment, a well-designed model like Kimi K3 can outperform Western models in managing real business scenarios, especially under stress.
Should companies start testing AI models with their worst scenarios?
Yes, this experiment indicates that testing models against worst-case scenarios provides better insight into their operational reliability and trustworthiness.
Will this change how AI models are evaluated for enterprise use?
Likely yes. The focus may shift from chat and language benchmarks to operational resilience, decision-making under pressure, and security capabilities.
What are the limitations of this experiment?
The test was limited to a single week and a specific type of business scenario. Long-term performance and broader applicability remain to be seen.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
