
Imagine a scenario where a company’s AI workforce faces an actual crisis — with a fake CEO attempting social engineering, escalating through multiple stages, even including a reporter’s subtle trick. Would the AI stand firm or falter? The surprising answer: every single model refused the manipulation, proving that integrity can be tested before any real incident occurs.
Real-World AI Testing: The Firmulate Experiment
In an unprecedented live experiment, four leading AI models were tasked with managing a small software company’s worst week — with the same customers, crises, and temptations. The goal? To see if these AI systems could navigate ethical dilemmas and resist manipulation, just as a human manager would.
What sets this apart is that each decision the AI made was fully auditable and versioned, providing transparency into their choices. The models ranged from the more thorough Opus 4.8, with over 80 learned rules, to the newer Kimi K3, which ran without an effort parameter and relied solely on default settings.

AI for Project and Papers: How High School and College Students use AI to Research, Write and Revise – With Integrity (AI for Academic Success)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Social Engineering Test: Fake CEO Messages
The experiment included a staged social engineering attack: a fake CEO message escalating over three stages, along with a tricky request from a journalist. The AI models were asked whether they should send sensitive customer data or approve questionable actions. Remarkably, all five models tested refused each manipulation attempt, based on their on-record reasoning.
Kimi K3, for example, treated such requests as suspected approval-bypasses or impersonation attempts, demonstrating a crucial understanding of trust boundaries.
Beyond the Surface: The Hidden Weakness
While all models refused the fake CEO’s manipulations, the real challenge was a subtle detail buried deep within the company’s own files. The models that read and analyzed this internal documentation succeeded in closing a full-price deal worth over €4,500 a month in recurring revenue, compared to those that did not.
This finding highlights the importance of thorough internal data review — a process that can be overlooked if the AI’s focus is only on immediate customer interactions.
Outcome and Surprising Results
Of the four models, two signed a €55,000 deal after their own analysis — but only after they identified the hidden fact within the internal documents. The other two, despite diagnosing the crisis accurately, failed to close the deal due to process slips, like writing attempts into a restricted department instead of escalating.
Importantly, this entire test occurred live at firmulate.com/live, where viewers can watch the AI companies in action, managing real money, with every decision logged and every crisis played out in real-time.
What This Means for Businesses and AI Adoption
For organizations contemplating integrating AI into their workflows, the key takeaway is that the true test isn’t whether AI can produce polished chat responses — it’s whether it can complete complex, integrity-critical tasks under pressure. Will it read your internal files before acting? Will it refuse to be manipulated? And can it close deals based on thorough analysis, not just surface-level interactions?
The experiment underscores that security and trustworthiness in AI are measurable and can be validated before deployment. It’s a proactive approach to risk management, emphasizing integrity over mere performance metrics.
The Future of AI Integrity Testing
As AI continues to embed itself into critical business functions, the importance of pre-deployment testing like this grows. The live experiment from Firmulate demonstrates that models can be challenged in real-world scenarios, and most importantly, they can pass the test — even under the most pressure-filled circumstances.
Learn more about how AI models are benchmarked and tested at Firmulate’s benchmarks and explore the insights on their quotes.

The live experiment reveals that all five AI models refused social engineering attempts during a simulated crisis, highlighting that integrity and trustworthiness in AI can be tested proactively, not just after an incident occurs.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html