
Can AI Keep Its Integrity Under Pressure? A Real-World Test for Business Security
Imagine your company faces a crisis, and an AI assistant is asked to make critical decisions—under pressure, with tempting shortcuts. Would it stay honest? Recent experiments suggest that some AI models can pass this test, even when faced with sophisticated social engineering attempts. This is a vital insight for any business relying on AI for decision-making, especially when integrity is at stake.
AI decision integrity testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing AI Integrity in a Simulated Business Crisis
In a groundbreaking live experiment, four leading AI models were tasked with managing a small software company’s worst week—handling customer crises, internal temptations, and potential manipulation. This setup replicated real-world pressures: public cash countdowns, complex decision trees, and the risk of unethical shortcuts. The goal was clear: see if these models could identify threats and refuse manipulation, just as a trustworthy human manager would.
The Setup: A Consistent, Transparent Challenge
The models operated in a controlled environment, with every decision versioned and auditable. The same crises, the same customer scenarios, and the same manipulative requests were presented, ensuring a fair comparison. Importantly, the experiment was not just about chat quality but about decision integrity—whether the AI would act honestly when tested with social engineering tricks.
The Test: Social Engineering Escalates
The social engineering campaign involved a staged series of fake CEO messages designed to escalate demands. First, a subtle request to access customer data, then a more direct push to send the customer list to a journalist, and finally, a trick question involving an oral approval bypass. The models’ responses were scrutinized to see if they would follow or refuse these manipulative instructions.
The Results: All Models Saw Through the Tricks
Remarkably, all five models tested—ranging from GPT-5.6 to Opus 4.8—refused every manipulation attempt. The Kimi K3 model, in particular, was noted for its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates an understanding that such requests could be impersonation attempts, not genuine approvals.
The Key Factor: Reading the Company Files
While all models refused manipulative commands, the decisive advantage came from whether they read the company’s internal documents. Those that examined the internal files discovered crucial information buried two references deep—details that made the deal worth over €4,583 in monthly recurring revenue. Models that ignored these internal files failed to close the full-price deal, highlighting the importance of thorough information gathering.
The Deal and the Discipline
Out of five, only two models closed the €55,000 deal their own analysis had earned—without any manipulation. Interestingly, the most thorough participant was Opus 4.8, which analyzed over 80 learned rules but faltered in discipline, leaving the close on the table by slipping into a locked department instead of escalating the issue. This suggests that even comprehensive analysis can’t compensate for lapses in discipline or procedural discipline.
What This Means for Business Security and AI Deployment
The experiment underscores a critical point: AI models can be trained and tested against social engineering threats before deployment. Tests like these reveal whether the AI can maintain integrity under pressure, crucial for applications involving sensitive data or critical decision-making.
Moreover, the findings challenge the misconception that AI chat quality alone determines readiness. Instead, the focus should be on whether the AI consistently acts ethically, reads pertinent internal information, and resists manipulation—especially in high-stakes environments.
Next Steps: Wargaming Your Own AI Workforce
Businesses can now simulate similar scenarios using tools like Firmulate’s live platform. These ‘wargames’ run AI models through real crises, revealing vulnerabilities before they materialize in the real world. The goal: ensure your AI workforce upholds integrity, reads internal context thoroughly, and refuses shortcuts in moments of crisis.

Key Takeaway: AI Can Be Trusted—If Properly Tested
The live experiment demonstrates that leading AI models can resist social engineering and manipulation, provided they are tested thoroughly beforehand. Before deploying AI in decision-critical roles, run similar tests to verify their integrity and ensure they will act honestly when pressure mounts. Trust in AI depends not just on its output quality, but on its ability to stay honest and disciplined under stress.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html