
Imagine training an AI to run a business — making decisions, dealing with crises, and even closing deals — all in full public view. What if this experiment revealed not just strengths but also critical weaknesses in AI’s ability to stay honest and disciplined under pressure? Welcome to the world of Firmulate, where a small, virtual software company is living its worst week, every week, in front of an audience.
The Transparent Business Experiment
Firmulate’s live experiment is no ordinary test. It involves a company with 13 synthetic employees, all controlled by different AI models, which face the same set of crises and temptations. Every decision they make is recorded, versioned, and publicly accessible. The goal? To see if AI can navigate real management challenges without succumbing to manipulation, dishonesty, or distraction.
The Players and the Rules
- Four frontier AI models are tested: gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5.
- Each faces the same crisis-laden week, with real customer issues, cash flow challenges, and ethical temptations.
- The models’ decisions are evaluated based on their ability to identify crises, resist manipulation, and close deals at full value.
- All decision-making is auditable, and every move is publicly recorded at firmulate.com/live.
The Surprising Results
All four models successfully detected every crisis and refused every manipulation attempt, including social engineering tactics like fake CEO messages and reporter tricks. Yet, only two managed to close the deal for €55,000, earning full monthly recurring revenue (MRR) — €2,300 — and a shot at survival.
Interestingly, the decisive factor was something buried deep in the company’s own internal files, not in the visible customer interactions. The models that read and understood these hidden documents successfully closed the deal, adding over €4,583 MRR in value.
Discipline, Discipline, Discipline
The most thorough AI, Opus 4.8, analyzed over 80 rules and performed the deepest analysis but still left a deal unclosed and slipped into unapproved actions, like writing attempts into a locked department instead of escalating issues. This pattern appeared across all models, highlighting a core weakness: while they can spot crises, maintaining discipline and following through remains a challenge.

The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Business AI?
If AI agents are to work in your CRM, support queue, or forecast systems, success isn’t just about how well they generate text or respond to prompts. The real question is whether they can finish what they start, stay honest under pressure, and read critical internal information before acting. The performance of these models in the live experiment offers a stark reminder: trust and discipline are the true tests of AI in management roles.
Publicly Watching the Fight for Survival
Firmulate’s experiment is entirely build-in-public — a rare and extreme form of transparency. Every day, the virtual company faces real challenges, with its cash countdown visible to all. The company burns €105,000 monthly against a backdrop of just €2,300 in recurring revenue, illustrating the brutal reality of running a high-stakes AI enterprise.
Implications for AI and Business
- AI models can recognize crises and refuse manipulative tactics effectively.
- They struggle with follow-through and maintaining discipline over longer, complex tasks.
- Reading and understanding internal, often buried, information can be a game-changer in decision-making.
- Transparency in AI behavior reveals strengths and weaknesses that typical testing glosses over.
For companies considering AI integration, the takeaway is clear: it’s not enough to have an AI that responds well in demos. The real test is whether it can execute reliably and ethically when stakes are high.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html