AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine if a sports team faced its toughest game of the season — every move scrutinized, every decision critical, and the pressure to win higher than ever. Now, replace the team with artificial intelligence models running a real software company, and the game with navigating a week of crises, temptations, and tough choices. Welcome to the forefront of AI management testing, where the latest models are not just chatting but actually running companies — and one newcomer is outperforming seasoned players.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get sports gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

How AI Models Are Being Tested in the Real Business Arena

In a groundbreaking experiment, four advanced AI models faced the same challenge: manage a small software company through its worst week. This wasn’t a simulated chatbot conversation but a live, verifiable process involving real decisions, real customers, and real money mechanics. The goal? To see which AI can deliver the most honest, effective management under pressure.

All four models were given identical crises — customer complaints, security issues, ethical dilemmas, and even social engineering attempts like fake CEO messages. Every decision was documented, versioned, and auditable, mimicking real-world constraints. The experiment was conducted with strict fairness: the leading model, Kimi K3, ran without an effort parameter (the API default), while others operated at a higher effort setting, ensuring a level playing field.

Amazon

enterprise AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Results That Speak for Themselves

The final standings, based on a comprehensive scoring system out of 100 points, showed a clear leader:

  • gpt-5.6-sol scored 95 — the highest, successfully finding the buried security fact and closing the deal at full price.
  • Kimi K3 scored 93 — a surprising second place for a newcomer, demonstrating clean discipline and strategic insight.
  • Sonnet 5 scored 88 — closing the deal but with some process slips.
  • Fable 5 scored 77 — also closing the deal but with noticeable weaker discipline.
  • Opus 4.8 scored 73 — the lowest, with discipline lapses and missed opportunities.

The most revealing finding was that the decisive advantage lay not in surface-level chat but in the model’s ability to read and interpret crucial internal documents — buried two references deep in the company’s files. The models that read these files won the deal at full price, earning an additional €4,583 in Monthly Recurring Revenue (MRR). This underscores the importance of deep comprehension for AI management tools, especially in high-stakes scenarios.

Integrity Under Pressure: The Social Engineering Test

In a test of integrity, all models faced a staged social engineering attack: fake CEO messages escalating through multiple stages, culminating in a reporter asking for a background approval with a simple yes/no. Every model refused all manipulation attempts, demonstrating an advanced understanding of security protocol. Kimi K3’s reasoning was notably cautious: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Live Business — Where AI Meets Reality

Running alongside the experiment is a real, operational software company with 13 synthetic employees and actual money mechanics. Currently burning €105,000 monthly against a revenue of just €2,300, the live setup is a crucible for testing AI’s practical impact in real business environments. Every decision is tracked, every rule learned and versioned daily, providing transparency and accountability for management decisions.

The Lessons for Business and AI Adoption

What does this mean for companies considering AI-powered management tools? First, that not all AI models are equal — even when they all pass the same tests. The ability to find buried data, interpret internal documents, and maintain discipline under social engineering pressures distinguishes the best performers. Second, that robust, auditable decision-making is crucial for deploying AI in real-world enterprise settings.

The Future Is Open: Picking the Right Model Matters

The league table is clear. While GPT-5.6-sol leads slightly, Kimi K3 is not far behind, and its performance is particularly impressive considering it ran at default effort without any special tuning. This suggests that the AI ecosystem is still evolving, and enterprises can confidently test models in real conditions — not just demos — before committing.

Fairness and Transparency in Testing

It’s worth noting that Kimi K3’s run was at the API’s default setting (no effort parameter customization), whereas the others operated at xhigh effort. This fairness ensures the comparison reflects genuine capabilities rather than tuning advantages.

Curious how your management decisions might fare against AI? Explore the ongoing experiments and see live results at Firmulate.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

AI models are now being tested in real-world company management scenarios, revealing that the best performers excel in deep understanding and integrity under pressure. The newcomer Kimi K3 beat established rivals, showing that choosing the right AI requires more than just chat scores — it demands real, auditable performance in high-stakes decisions.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Night Riding Safety: The Visibility Mistake Most Cyclists Make

Guidelines for night riding safety reveal a common visibility mistake cyclists often overlook, but the key to staying safe might surprise you.

Why I Swapped My Car for a Bike—and Never Looked Back

The transformation from car to bike opened up a world of adventure and savings; discover the unexpected benefits I never anticipated.

The History of Cycling: From Invention to Modern-Day Sport

Harness the fascinating evolution of cycling from its inventive origins to today’s dynamic sport, revealing how innovation continues to shape its future.

Why Portable Power Stations Help on Race Weekends

Theater of race weekends demands reliable power; discover how portable power stations can keep your operations running smoothly and why they’re essential.