
Imagine pushing your fitness routine to its limit — not just lifting weights but navigating real setbacks, temptations, and tough choices. Now, replace that with managing a business during its worst week. That’s exactly what a groundbreaking live experiment is doing with AI models, and the results could change how we think about automation and management.
The Reality of Business Management Goes Beyond Chatbots
Many AI demonstrations focus on impressive chat responses, but real management is about much more than answering questions. It’s about making tough decisions under pressure, reading critical files, maintaining honesty, and completing complex tasks — even when temptations to cut corners are high. A recent public experiment, run at Firmulate, brings this reality into focus.
As an affiliate, we earn on qualifying purchases.
The Live Business Wargame
In this unique test, four frontier AI models were tasked with running a small software company through its worst week. The scenario included the same customers, crises, and temptations for each model, with every decision documented and auditable. The goal? See who could handle the pressure, identify critical facts buried deep in the company’s files, and keep honest — all while closing business deals.
Key Findings: More Than Just Answers
- All models detected every crisis and refused every manipulation attempt — a promising start for trustworthiness.
- Only two models signed the €55,000 deal their analysis earned. The others identified the right opportunities but failed to close, leaving money on the table.
- The decisive factor was not superficial answer quality but the ability to uncover a key fact buried two document references deep in the company’s own files. The models that read the files and used that information won the deal at full price, adding +€4,583 MRR.
- The models refused social engineering attempts, including fake CEO messages and a reporter trick, demonstrating integrity under pressure.
- The live operation involved 13 synthetic employees, real money mechanics burning €105K/month against €2.3K MRR, and a suite of over 680 self-learned rules guiding daily decisions. You can watch this unfold every business day at firmulate.com/live.
Beyond Chat: The Management Test
This experiment exposes a crucial truth: performance in chat demos doesn’t equate to effective management. The real question is whether an AI can finish what it starts, read critical information, stay honest under pressure, and deliver measurable work — not just generate convincing responses.
The League Table and What It Means
| Model | Score | Verdict |
|---|---|---|
| gpt-5.6-sol | 95 | Found the buried fact, closed the deal — the complete performance. |
| Kimi K3 | 93 | Closed the deal too, with the cleanest discipline of the field. |
| Sonnet 5 | 88 | Closed the deal with minor slips. |
| Fable 5 | 77 | Closed the deal but with some process slips. |
| Opus 4.8 | 73 | Left the close on the table, discipline slipped. |
This isn’t just about AI chat quality but about evaluating what management quality truly looks like — honesty, thoroughness, and resilience under stress. As these models get better, the real test is whether they can handle the complex, high-stakes decisions of a real business, not just generate plausible responses.
Why You Should Care
If AI tools will manage your CRM, support queues, or forecast future needs, then the key question becomes: can they finish what they start under pressure? Do they read your files first? Do they stay honest? And crucially, what is the cost per unit of useful work produced? The experiment from Firmulate shows us that simply passing chat benchmarks isn’t enough — management skills matter more.
Learn More and Watch Live
Curious about how different models perform? Visit firmulate.com/benchmarks.html for full results and plain-language insights. Want to test your own enterprise? You can run the same wargame against your business data — nothing writes back to your systems, but it reveals how your AI workforce might handle genuine crises. Find out more at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html