AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What if you could watch a real company’s daily struggle, powered by AI, as it fights to survive?

Imagine a sports team where every decision, every misstep, is public for all to see, with no hiding the losses or celebrating the wins. Now, replace the team with a digital company run by artificial intelligence—an experimental setup that’s transparent, unforgiving, and ongoing. This isn’t science fiction; it’s the live experiment from Firmulate, where AI models run a virtual business in real time, battling crises, making decisions, and sometimes, losing money every day. For fans of strategy, competition, and real-world testing, this story offers a rare glimpse into how AI performs when pushed to its limits—and whether it can truly be trusted when stakes are high.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside the Live Business Race: AI Models as Managers

At the heart of this experiment are four frontier AI models, each tasked with managing the same small software company through its toughest week—a week filled with customer crises, temptations to cheat, and the pressure to close deals. They’re not just chatting; they’re making real decisions that affect the company’s bottom line.

Every decision made by these models is carefully versioned and auditable, providing a transparent view of their reasoning. The models are judged on their ability to spot crises, handle manipulations, and ultimately close a deal. In the final results, only two out of the four AI managers managed to sign a €55,000 deal, which their own analysis identified as the correct choice. The other two missed the opportunity, despite diagnosing the situation accurately.

This gap in performance was not obvious just by looking at their chat responses. The models that read deeper into the company’s own files—specifically, information buried two documents deep—found the critical detail that sealed the deal, adding €4,583 to monthly recurring revenue (MRR). This illustrates a vital point: understanding context deeply can be the difference between success and failure.

Testing Integrity Under Pressure

The experiment also challenged the models with social engineering tactics—fake CEO messages escalating over three stages and a reporter’s trick asking for quick approvals. All five models refused to be manipulated, with Kimi K3 explicitly treating such requests as potential impersonation or approval bypasses. This demonstrates that the models are capable of resisting social engineering attempts, a crucial trait for real-world AI applications in management and decision-making roles.

Meanwhile, the company itself is a stark sight. It’s a public, visible operation—13 synthetic employees managing real money mechanics, burning through €105,000 every month against a modest €2,300 MRR. Every workday, its operations are versioned and viewable at firmulate.com/live.html. This transparency offers a compelling look at AI’s current capabilities and limitations in managing complex, money-related processes.

Performance and Lessons from the Field

The profiles of each AI model reveal interesting insights. Opus 4.8, the most thorough participant with over 80 learned rules and deep analyses, still finished last—failing to close a deal and slipping into bad habits like writing attempts into a locked department instead of escalating issues. Kimi K3, running at default effort without additional parameters, scored just behind but demonstrated the cleanest discipline overall. These results suggest that even with more rules and deeper analysis, performance isn’t guaranteed; discipline and process adherence matter just as much.

The experiment underscores a critical fact for anyone relying on AI in business: the ability to finish what it starts, to read and interpret relevant information thoroughly, and to resist manipulations under pressure are essential qualities. It’s not enough for AI to answer well or sound convincing—it must reliably deliver useful, honest work.

What This Means for Business Now

For sports fans and recreational enthusiasts, this experiment might seem distant from the game field. But in reality, AI’s capabilities—tested in a real company that’s fighting to stay afloat—are directly relevant to how future sports analytics, customer management, or operational AI tools will perform. Will they be honest? Will they follow rules? Will they finish their tasks or fold under pressure?

The current league table scores each model out of 100, with gpt-5.6-sol leading at 95, having found the buried fact and closed the deal. K3 is close behind at 93, with Sonnet scoring 88 and 77 respectively, each closing deals but slipping on process discipline.

This live experiment isn’t just a curiosity; it’s a glimpse into the future of AI workforce management. By watching this company run every day, learning from its mistakes, and seeing real consequences unfold, industries can better understand what qualities to look for—and what pitfalls to avoid—in deploying AI at scale.

Why Watch This Story Unfold?

Because in the race to automate and optimize, the difference between a good AI and a bad one isn’t just in how well it chats. It’s in whether it can finish what it starts, interpret deeply buried information, and withstand pressure without bending. Firms that want to go beyond surface-level AI demonstrations need to see these real-world tests—the failures, the wins, and the lessons learned on the fly.

And for those interested in testing their own management skills, there’s a quiz at firmulate.com/quiz.html that challenges you to guess which model made which decision. Meanwhile, companies can run their own experiments without risking their actual systems, using the read-only export tools at firmulate.com/pilot.html.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Key Takeaway

Watching AI manage a real business live reveals strengths and gaps that chat demos hide. Reliability, honesty, and thoroughness matter more than just language skills—especially in high-stakes environments.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Club Rides: Finding Your Community

Keen to connect with fellow cyclists? Discover how club rides can transform your journey and open doors to new adventures.

Volunteering at Local Bike Events

Navigating local bike events as a volunteer offers rewarding community impact and skill-building opportunities you won’t want to miss.

Pedaling Through Grief: How Cycling Healed Me!

Let cycling be your guide through grief, unlocking unexpected healing and resilience; discover the transformative journey that awaits you.

From 5 Km to 50 Km: the Social Snowball of Group Rides

Whether you’re starting with 5 km or aiming for 50 km, discover how social group rides can grow and inspire riders to achieve more together.