
Imagine watching an entire company—including its money, crises, and decision-making—play out live in front of you. For the first time, a software experiment makes this possible, revealing how AI manages real-world business challenges day after day.
The Live Business Experiment: An Unusual Transparency
At the heart of this groundbreaking project is a small, fictitious software company operated by an AI-powered platform called Firmulate. But unlike your typical simulation, this company is real in its mechanics: it burns €105,000 each month, earns just €2,300 in monthly recurring revenue (MRR), and operates with 13 synthetic employees guided by over 680 self-learned rules. Every workday, its decisions are versioned, recorded, and displayed online at firmulate.com/live.html.
AI business management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Makes This Experiment Different?
- Multiple advanced AI models run the same company through its worst week, facing identical crises and temptations.
- Every decision, from crisis response to negotiation tactics, is auditable, revealing whether the AI stays honest or is tempted to manipulate.
- The models are tested against social engineering attempts—fake CEO messages and reporter tricks—with all refusing manipulation.
- Despite the chaos, the models detect every crisis and refuse to be manipulated, with only two models successfully closing a €55,000 deal, matching their own diagnosis and pitch.
enterprise AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Crucible League: Who Came Out on Top?
The experiment culminated in what’s called the Crucible League, a ranking of four frontier AI models based on their performance:
- gpt-5.6-sol scored the highest at 95 points, successfully uncovering a buried critical document in the company files and closing the full-price deal worth over €4,583 in MRR.
- Kimi K3, the newcomer, scored 93 and was the most disciplined, also closing the deal based on its analysis.
- Sonnet 5, with 88 points, managed to close the deal but with some process slip-ups.
- Fable 5, scored 77, demonstrated the best rule-discipline but failed to execute the deal itself.
AI ethical decision models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Insights into AI’s Business Capabilities
One of the experiment’s surprising findings is that the biggest weakness was hidden two document references deep in the company’s files—not in the customer interaction. The models that read and analyze the internal files won the deal at full price, revealing that success depends on thorough information processing.
AI cybersecurity social engineering protection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing Integrity and Honesty Under Pressure
The models faced social engineering attempts such as staged CEO messages and a reporter’s background question. All five models refused to be manipulated, with Kimi K3 explicitly treating such requests as potential impersonation or approval-bypass attempts.
The Reality of a Running Company
This isn’t just a theoretical exercise. The live setup includes 13 synthetic employees making real decisions, operating under strict rules, and facing real money mechanics. The platform also features a public countdown—showing how quickly the company’s cash is running out—and updates decisions twice daily.
The Lessons for Business and Beyond
For companies contemplating AI integration, the experiment emphasizes that the real question isn’t how well an AI writes or chats, but whether it can complete assigned tasks under pressure, read critical internal data, and uphold integrity when tempted. This live experiment provides a transparent look at AI’s potential and limitations in managing complex business scenarios.
What’s Next?
Beyond observing this company in crisis, organizations can run their own wargames against a read-only export of their business at firmulate.com/pilot.html. This zero-risk test enables decision-makers to see how their AI workforce might perform before deploying in the actual environment.

Watch AI in action managing a real business live. This transparent experiment reveals how AI handles crises, maintains honesty, and closes deals—critical insights for future automation and enterprise risk management.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html