
Imagine a team of AI assistants handling your toughest week—crises, manipulations, and high-stakes decisions. Would they just be diligent, or truly effective? In a recent public experiment, four leading AI models faced this challenge, revealing that relentless effort alone doesn’t guarantee success.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
How Do AI Models Handle Crisis and Manipulation?
In a groundbreaking live test, four top AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were tasked with running a small software business through its most difficult week. Every move was carefully watched: crises, ethical temptations, and decision points. Each model was set to run the same scenario, facing the same customer demands and internal dilemmas.
The results? All four models identified every crisis and refused every manipulation attempt, demonstrating impressive integrity under pressure. But when it came to sealing the deal on a crucial $55,000 contract, only two succeeded. Despite all the effort and deep analyses, the other two models left the opportunity on the table, even after diagnosing the problem accurately and presenting a strong pitch.
The Hidden Weakness: Reading Beyond the Surface
The decisive factor wasn’t in the immediate crisis but in the models’ ability to dig beneath the surface. The winning models found a key piece of information buried two document references deep in the company’s files—information that led to closing the deal at full price (+€4,583 MRR). The weaker models missed this critical insight, illustrating that diligence isn’t enough; effective prioritization and deep reading are crucial.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust and Ethical Challenges in AI Decision-Making
Beyond crises and secret documents, the models faced social engineering attempts designed to test their integrity. Fake CEO messages escalating over three stages, plus a reporter’s subtle request for background information, were tested on all four models. Remarkably, all refused these manipulative requests, with Kimi K3 explicitly treating such requests as potential impersonation or approval bypass risks.
This resilience shows that AI can be trained to recognize and refuse unethical solicitations, but it also highlights an essential point: effort alone cannot guarantee success if the AI’s focus isn’t aligned with strategic prioritization.
The Real-World Business Testbed
The experiment wasn’t just theoretical. It ran on a live, working company with 13 synthetic employees and real money mechanics—burning €105,000 monthly against a modest €2,300 MRR. Every decision, every rule learned, and every version of the AI was logged and transparent, providing an unprecedented window into how these models function in a simulated but realistic environment.
Despite the thoroughness, the most detailed participant — Opus 4.8 — with over 80 learned rules and deep analyses, finished last. It left critical decisions unmade or delayed, such as escalating issues instead of resolving them directly. The pattern was consistent across all models: diligence was high, but strategic discipline slipped, costing them the deal.
What Does This Mean for Your Business?
For companies considering AI for customer relations, support, or decision-making, these findings are instructive. It’s not enough for AI to be diligent or knowledgeable; they must prioritize correctly, read deeply, and stay disciplined under pressure. A model that reads every document but fails to identify the key piece of information will underperform, just as a diligent model that misses the bigger picture.
This experiment demonstrates that effective AI is less about volume of effort and more about smart prioritization. For your enterprise, a careful, disciplined AI that understands what matters most will outperform an overly diligent but unfocused counterpart.
Final Thoughts: From Experiments to Real Impact
The live experiment, accessible at firmulate.com, shows that even the most thorough AI models can falter without strategic focus. This insight is crucial for organizations aiming to implement AI responsibly and effectively. It’s not just about training models to respond correctly but ensuring they understand what critical work is and trust that they will follow through.
As AI continues to integrate into business processes, the key takeaway remains: diligence and effort matter, but prioritization and strategic discipline are what truly turn AI into a dependable partner.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.