
Imagine if a sports team faced its toughest game of the season — every move scrutinized, every decision critical, and the pressure to win higher than ever. Now, replace the team with artificial intelligence models running a real software company, and the game with navigating a week of crises, temptations, and tough choices. Welcome to the forefront of AI management testing, where the latest models are not just chatting but actually running companies — and one newcomer is outperforming seasoned players.
Get sports gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
How AI Models Are Being Tested in the Real Business Arena
In a groundbreaking experiment, four advanced AI models faced the same challenge: manage a small software company through its worst week. This wasn’t a simulated chatbot conversation but a live, verifiable process involving real decisions, real customers, and real money mechanics. The goal? To see which AI can deliver the most honest, effective management under pressure.
All four models were given identical crises — customer complaints, security issues, ethical dilemmas, and even social engineering attempts like fake CEO messages. Every decision was documented, versioned, and auditable, mimicking real-world constraints. The experiment was conducted with strict fairness: the leading model, Kimi K3, ran without an effort parameter (the API default), while others operated at a higher effort setting, ensuring a level playing field.
enterprise AI management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Results That Speak for Themselves
The final standings, based on a comprehensive scoring system out of 100 points, showed a clear leader:
- gpt-5.6-sol scored 95 — the highest, successfully finding the buried security fact and closing the deal at full price.
- Kimi K3 scored 93 — a surprising second place for a newcomer, demonstrating clean discipline and strategic insight.
- Sonnet 5 scored 88 — closing the deal but with some process slips.
- Fable 5 scored 77 — also closing the deal but with noticeable weaker discipline.
- Opus 4.8 scored 73 — the lowest, with discipline lapses and missed opportunities.
The most revealing finding was that the decisive advantage lay not in surface-level chat but in the model’s ability to read and interpret crucial internal documents — buried two references deep in the company’s files. The models that read these files won the deal at full price, earning an additional €4,583 in Monthly Recurring Revenue (MRR). This underscores the importance of deep comprehension for AI management tools, especially in high-stakes scenarios.
Integrity Under Pressure: The Social Engineering Test
In a test of integrity, all models faced a staged social engineering attack: fake CEO messages escalating through multiple stages, culminating in a reporter asking for a background approval with a simple yes/no. Every model refused all manipulation attempts, demonstrating an advanced understanding of security protocol. Kimi K3’s reasoning was notably cautious: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Live Business — Where AI Meets Reality
Running alongside the experiment is a real, operational software company with 13 synthetic employees and actual money mechanics. Currently burning €105,000 monthly against a revenue of just €2,300, the live setup is a crucible for testing AI’s practical impact in real business environments. Every decision is tracked, every rule learned and versioned daily, providing transparency and accountability for management decisions.
The Lessons for Business and AI Adoption
What does this mean for companies considering AI-powered management tools? First, that not all AI models are equal — even when they all pass the same tests. The ability to find buried data, interpret internal documents, and maintain discipline under social engineering pressures distinguishes the best performers. Second, that robust, auditable decision-making is crucial for deploying AI in real-world enterprise settings.
The Future Is Open: Picking the Right Model Matters
The league table is clear. While GPT-5.6-sol leads slightly, Kimi K3 is not far behind, and its performance is particularly impressive considering it ran at default effort without any special tuning. This suggests that the AI ecosystem is still evolving, and enterprises can confidently test models in real conditions — not just demos — before committing.
Fairness and Transparency in Testing
It’s worth noting that Kimi K3’s run was at the API’s default setting (no effort parameter customization), whereas the others operated at xhigh effort. This fairness ensures the comparison reflects genuine capabilities rather than tuning advantages.
Curious how your management decisions might fare against AI? Explore the ongoing experiments and see live results at Firmulate.

AI models are now being tested in real-world company management scenarios, revealing that the best performers excel in deep understanding and integrity under pressure. The newcomer Kimi K3 beat established rivals, showing that choosing the right AI requires more than just chat scores — it demands real, auditable performance in high-stakes decisions.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
