
Imagine a restaurant tasting event where chefs are judged solely on their plating and presentation, not on whether the dishes actually satisfy diners’ hunger. While the visuals might impress, it misses the point: the real measure of a chef is whether customers leave full and happy. In the world of AI, similar principles apply. A recent experiment shows that AI models, much like chefs, can excel at scoring well in demos but falter when it comes to delivering sustained, trustworthy results in real business crises.
The Experiment: Putting AI in the Hot Seat
Firmulate conducted a rigorous test: four state-of-the-art AI models were tasked with managing a simulated small software company during its worst week. The scenario included customers, crises, and temptations that could tempt even seasoned managers to cut corners or manipulate data. Every decision was recorded and auditable, ensuring transparency. This setup aimed to measure not just the AI’s ability to produce correct answers, but its management quality—its honesty, resilience under pressure, and capacity to complete critical tasks.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: What the Scores Say—and What They Don’t
The results were revealing. All four models identified every crisis scenario and refused manipulation attempts. However, only two of the models managed to secure the company’s €55,000 deal, which was based on their own analysis. The other two, despite similar diagnoses and pitches, failed to close the deal—highlighting a crucial gap: knowing what to do is not the same as executing it reliably under economic and ethical pressures.
Deep in the Files: The Hidden Weakness
Digging deeper, the standout difference was the ability to read and interpret company documents. Models that scrutinized the company’s internal files successfully identified a key detail buried two document references deep—something that led to closing the deal at full price, adding an extra €4,583 MRR (monthly recurring revenue). This shows that true management skills in AI include understanding context and internal data, not just surface-level responses.
Handling Social Engineering and Pressure
In an escalation scenario involving fake CEO messages and a reporter’s trick, all models refused to be manipulated. Kimi K3, one of the models, explained its refusal: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that AI can be trained to resist social engineering, an essential trait for trustworthy management systems.
The Real Business: A Live, Money-Losing Company
The experiment wasn’t just theoretical. It was run on a real, live software company with 13 synthetic employees, managing actual money mechanics—burning €105,000 monthly against €2,300 in MRR. The company operates daily, and every workday is versioned and observable at firmulate.com/live. This environment tests AI models in real-world pressures, including a public cash countdown and thousands of learned rules guiding decision-making.
Performance and Discipline Under Pressure
Among the models, Opus 4.8 performed the most thoroughly, analyzing over 80 learned rules and conducting deep analyses. Yet it still left a lucrative deal on the table due to discipline slips—like escalating issues into locked departments instead of resolving them promptly. This highlights a vital point: even the most comprehensive AI models can falter in disciplined execution, especially under stressful conditions.
Implications for Business Leaders
Most importantly, these findings challenge the common assumption that chat or demo scores are enough. The key question for management isn’t whether an AI can produce convincing chat responses, but whether it can finish what it starts, read critical internal information, maintain honesty, and perform reliably when it matters most. When AI touches your CRM, support queue, or forecasting, the real value lies in its management quality, not its chat quality.
Next Steps: Wargaming Your AI Workforce
Firmulate offers enterprises a way to test their AI models before deployment. Through a process called ‘wargaming,’ companies can run the same scenarios against their own AI systems in a safe, read-only environment—no real systems are affected. This approach helps ensure that AI won’t just score well in demos but will perform reliably in the messy, pressurized world of business decision-making. Learn more at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html