
Imagine trying to get fit by merely logging every workout without focusing on the quality or importance of each session. Despite diligent record-keeping, progress stalls. Similarly, in AI-driven business management, volume and thoroughness alone don’t guarantee success. The latest live experiment from Firmulate reveals why prioritization, discipline, and strategic focus are crucial—even for AI models tasked with managing complex companies.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Firmulate Experiment: An AI Test of Diligence vs. Impact
At the heart of this real-world trial was a straightforward question: Can AI models, given the same challenging scenario, outperform each other based solely on their diligence and analytical depth? Four frontier models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—each faced identical crises within a simulated small software company. The goal? Navigate a tough week filled with customer issues, internal dilemmas, and manipulations, then secure a €55,000 deal.
The Setup and the Stakes
All models ran the same workflow, decision points, and crisis scenarios, with every choice recorded and auditable. The company’s scenario involved real-like mechanics: 13 synthetic employees, daily decision-making, and an ongoing cash burn rate of €105,000/month against a revenue of only €2,300 monthly recurring revenue (MRR). The experiment’s data was openly accessible and live—anyone could observe the models’ decision-making process at firmulate.com/live.
The Results: Diligence Meets Reality
- The top scorer, gpt-5.6-sol, scored 95 out of 100, spotting every crisis, refusing manipulation, and ultimately closing the deal.
- Kimi K3, the newcomer, scored 93 and also closed the deal, demonstrating the cleanest discipline among the models.
- Sonnet 5 scored 88 and closed, but with slightly more process slips.
- Opus 4.8, despite being the most thorough with over 80 learned rules and deep analysis, scored only 73 and failed to close the deal.
- The baseline—doing nothing—scored 26, underlining the importance of proactive management.
Interestingly, all models identified every crisis and rejected manipulative or deceptive tactics, such as fake CEO messages or reporter tricks. Only two models, gpt-5.6-sol and Kimi K3, managed to close the deal based on their own analysis. The critical disadvantage for Opus 4.8? Its discipline slipped during the final stages—decisions meant to escalate issues into the appropriate department were instead left in locked documents, affecting outcome.
The Hidden Weakness: Reading the Company Files
The decisive advantage went to models that read and understood deeper in the company’s documentation. Those that accessed information buried two document references deep in the internal files secured the full-price deal, worth an additional €4,583 in monthly recurring revenue. This underscores a vital insight: surface-level diligence isn’t enough. Impact hinges on prioritization—reading deeply, understanding context, and acting decisively on crucial information.
As an affiliate, we earn on qualifying purchases.
Implications for Business AI and Management
This experiment challenges the common misconception that volume and thoroughness alone produce better results. Diligence must be coupled with strategic prioritization, discipline, and context awareness. In human terms, it’s like training hard but ignoring the most critical exercises; in AI, it’s about reading the right information and acting on it under pressure.
In this simulated environment, models that prioritized impactful information and maintained discipline in decision protocols were more successful. Conversely, the most thorough, Opus 4.8, faltered because it failed to escalate critical issues when it mattered most. This isn’t just about AI performance; it’s a mirror for real-world management—no matter how diligent or analytical you are, success ultimately depends on focusing on what moves the needle.
What Does This Mean for Your Business?
If AI is to support your CRM, support queues, or forecasting, the question isn’t just whether it writes well or analyzes deeply. It’s whether it can finish tasks, read relevant information thoroughly, stay honest under pressure, and make impactful decisions. The effectiveness of an AI workforce is measured by its ability to prioritize impact over volume.
For companies considering deploying AI models, the message is clear: test your AI in a wargame before trusting it with your real business. Firms like Firmulate offer live, transparent experiments that emulate your company’s worst week—giving you insight into how your AI will perform in tough conditions.
Final Takeaway: Diligence Isn’t Enough—Prioritization Is Key
The live experiment vividly illustrates that no matter how thorough or learned an AI model is, its success depends on its ability to prioritize critical information and maintain discipline. The most effective models weren’t just the smartest—they were the most strategic, recognizing what truly mattered and acting decisively.
In the end, this isn’t just about AI. It’s a lesson for all managers and decision-makers: focus on impact, prioritize ruthlessly, and discipline your teams and tools to act on what counts the most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.