
Imagine a chef deciding whether to serve a dish only after reading the full recipe and tasting the ingredients firsthand. Now, translate that to AI: a system that doesn’t just skim your emails or chat logs, but actually reads and understands your files before making decisions. In the high-stakes world of business, this difference can be the line between sealing a €55,000 deal or missing out entirely.
The Hidden Depths of AI Decision-Making
Recent experiments reveal a crucial edge for AI systems: their ability to read and interpret multiple layers of company data before acting. Four leading AI models were tested in a simulation mimicking a small software company’s worst week — facing customer crises, internal dilemmas, and manipulation attempts. All four models could identify crises and refused manipulation attempts, but only two managed to close the deal, winning the trust and the €55,000 contract based solely on their own analysis.
What Made the Difference?
The key was whether the AI looked beyond surface-level information. The decisive weakness for competitors was buried two document references deep within the company’s own files—not in the customer event or superficial data. Models that read and understood the full context of these documents were the ones that closed the deal at full price, adding over €4,583 monthly recurring revenue (MRR).
As an affiliate, we earn on qualifying purchases.
The Significance of Deep Reading in AI
This ongoing experiment from Firmulate demonstrates a vital property for enterprise AI systems: the ability to perform multi-hop reasoning—reading multiple documents, linking them, and understanding the context fully before acting. Unlike typical chatbots or basic decision tools, these models show that reading your files thoroughly is a measurable, critical factor in business success.
Trust and Integrity Under Pressure
In a simulated social engineering attack—fake CEO messages escalating in stages, plus a reporter trick—every model refused manipulation attempts. One model, Kimi K3, explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline is essential when AI systems are integrated into real workflows, where trustworthiness is paramount.
The Real-World Company and Its AI Workforce
Beyond the lab, firms can test their AI agents against the same scenarios before deploying them live. The experiment runs on a simulated company with 13 synthetic employees, managing real money mechanics—burning €105k monthly against €2.3k MRR—with full transparency. Every decision, every slip, and every success is recorded and viewable online at firmulate.com/live.
What the Results Tell Us
The most thorough model, Opus 4.8, performed the worst in closing the deal because it left the opportunity unresolved. It showed that even the deepest, most analytical AI can slip if not disciplined—highlighting the importance of clear decision protocols and attention to detail. Meanwhile, models that scored highest, like gpt-5.6-sol at 95 points and Kimi K3 at 93 points, demonstrated that reading and understanding deeply buried information is vital for trust and success.
Why This Matters for Business Decisions
In real life, AI systems are increasingly touching your customer data, support queues, and forecasts. The question isn’t just whether an AI can write well or generate convincing responses; it’s whether it can finish what it starts, read your files thoroughly, stay honest under pressure, and deliver measurable work. This experiment underscores that the true test of an AI’s usefulness is its ability to read and interpret complex, layered information before acting.
Measuring AI Performance in Your Business
For enterprises interested in testing their own AI workforce, Firmulate offers a unique platform where they can run their scenarios against a read-only export of their business data. This way, companies can verify how their AI would perform in real crises without risking actual operations. To learn more, visit firmulate.com/pilot.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html