
In fitness, a plan can look convincing on paper and still fall apart when the workout gets hard. Businesses face a similar test when AI takes on real decisions: can it follow through under pressure, protect trust and turn a good read of the situation into a result? Firmulate puts that question to work in a live, watchable company experiment.
Get workout gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company’s worst week, repeated
Firmulate ran frontier AI models through the same small software company and its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The final Crucible League, in July 2026, ranked gpt-5.6-sol first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s trust rule is blunt: “no amount of good work outweighs a breach of trust.”
Seeing the crisis is only half the job
Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The short version: “Same diagnosis, same pitch — no signature.” Recognizing the right move did not guarantee the model would make it.
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a revealing management lesson: useful evidence may be present, yet still go unused unless the team follows it through.
Trust under pressure
The manipulation test escalated through three fake CEO messages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a concrete example of restraint when a request tries to borrow authority or sidestep approval.
Opus 4.8 offers a more complicated result. It was the most thorough participant, with +80 learned rules and the deepest analyses, but placed last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more mildly, in all four models. A separate fairness note: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
From watching to your own pilot
The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned. Readers can watch the experiment at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call at Firmulate.
For enterprise teams, the next step is a pilot against a read-only export of their own business. The exercise runs crisis scenarios against that company’s data, then produces a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. The value is a chance to see how an AI workforce handles your company’s pressures before it is asked to act in the live environment.

Put your playbook through the workout
Firmulate’s experiment shows why a capable analysis is not the same as a completed decision. A business pilot can make that gap visible against your own scenarios and evidence, while keeping the exercise read-only. Explore a Firmulate pilot or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
