
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
What Can a ‘Do-Nothing’ AI Score Tell Us About Business Readiness?
Imagine you’re hiring an employee who promises to handle your company’s toughest week. Now, what if that candidate barely tries but still scores 26 out of 100? Surprisingly, in the world of AI benchmarking, such a modest score reveals much about the system’s honesty and reliability — qualities critical whether you’re managing a restaurant or a software startup.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Truth Behind the Baseline Score
In a recent public experiment conducted by Firmulate, four advanced AI models faced identical challenges: managing a small software company’s worst week. Every decision was scrutinized, every crisis monitored, and every temptation to cheat observed. The results offer a revealing lens into the nature of AI performance metrics.
The ‘do-nothing’ baseline — essentially an AI that doesn’t try to improve or act proactively — scored 26 points. This isn’t zero because even minimal effort, like avoiding obvious mistakes or not falling for manipulative tactics, earns partial credit. It’s a reminder that in real-world applications, some performance is better than none, but it’s often not enough.
Partial Progress and the Trust Cap
One key insight from the experiment is that partial success counts. All models identified every crisis, refused manipulative tricks, and maintained honesty during social engineering tests. Yet, only two managed to close the deal on their own analysis and sign the €55,000 contract. The others, despite good diagnosis, failed to follow through or sign — illustrating that progress alone isn’t enough. A single breach of trust caps the total score, emphasizing that integrity matters as much as competence.
What Makes a Model Truly Ready?
Interestingly, the models that read deeply into a company’s files — not just surface-level data — were the ones that won the contracts at full price, worth over €4,500 in monthly recurring revenue. This highlights a crucial lesson: understanding context and verifying information thoroughly distinguishes trustworthy AI from superficial ones.
Tested Under Pressure: Honest Decisions Count
Throughout the experiment, all models refused social engineering attempts, such as fake CEO messages and reporter tricks, reinforcing their reliability. Kimi K3’s on-record reasoning, for instance, was to treat suspicious requests as potential impersonation, showing an explicit approach to maintaining integrity under pressure.
The Real-World Company in Action
The experiment wasn’t just theoretical. It involved a live simulation of a company with 13 synthetic employees, real cash flow mechanics, and a public cash countdown. Every decision the AI made was logged and auditable, making the entire process transparent and observable at firmulate.com/live. This setup reveals how AI behaves in real business settings, not just in polished demos.
Why the Gap Matters — and What It Tells Us
The AI that scored highest, gpt-5.6-sol, achieved a perfect 95, thanks to its ability to uncover the buried facts and close the deal. Kimi K3 followed closely at 93, demonstrating the importance of discipline and thoroughness. Meanwhile, Opus 4.8 and Sonnet scored 73 and 88 respectively, showing that even with many learned rules, some discipline slips can lead to missed opportunities.
Implications for Business and AI Adoption
For companies considering AI tools, these findings underline a vital point: performance isn’t just about language fluency or seeming cleverness. It’s about trustworthiness, consistency, and the ability to deliver results under pressure. An AI that refuses manipulation and reads the full context can be more valuable than one that just sounds convincing in a chat.
More broadly, the experiment demonstrates why honest benchmarks matter. They give a transparent, real-world view of AI readiness, helping decision-makers avoid overestimating their tools based on superficial metrics. As firms consider deploying AI in critical functions — from customer management to financial decisions — these benchmarks offer an honest yardstick.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
