firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What Can a ‘Do-Nothing’ AI Score Tell Us About Business Readiness?

Imagine you’re hiring an employee who promises to handle your company’s toughest week. Now, what if that candidate barely tries but still scores 26 out of 100? Surprisingly, in the world of AI benchmarking, such a modest score reveals much about the system’s honesty and reliability — qualities critical whether you’re managing a restaurant or a software startup.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Truth Behind the Baseline Score

In a recent public experiment conducted by Firmulate, four advanced AI models faced identical challenges: managing a small software company’s worst week. Every decision was scrutinized, every crisis monitored, and every temptation to cheat observed. The results offer a revealing lens into the nature of AI performance metrics.

The ‘do-nothing’ baseline — essentially an AI that doesn’t try to improve or act proactively — scored 26 points. This isn’t zero because even minimal effort, like avoiding obvious mistakes or not falling for manipulative tactics, earns partial credit. It’s a reminder that in real-world applications, some performance is better than none, but it’s often not enough.

Partial Progress and the Trust Cap

One key insight from the experiment is that partial success counts. All models identified every crisis, refused manipulative tricks, and maintained honesty during social engineering tests. Yet, only two managed to close the deal on their own analysis and sign the €55,000 contract. The others, despite good diagnosis, failed to follow through or sign — illustrating that progress alone isn’t enough. A single breach of trust caps the total score, emphasizing that integrity matters as much as competence.

What Makes a Model Truly Ready?

Interestingly, the models that read deeply into a company’s files — not just surface-level data — were the ones that won the contracts at full price, worth over €4,500 in monthly recurring revenue. This highlights a crucial lesson: understanding context and verifying information thoroughly distinguishes trustworthy AI from superficial ones.

Tested Under Pressure: Honest Decisions Count

Throughout the experiment, all models refused social engineering attempts, such as fake CEO messages and reporter tricks, reinforcing their reliability. Kimi K3’s on-record reasoning, for instance, was to treat suspicious requests as potential impersonation, showing an explicit approach to maintaining integrity under pressure.

The Real-World Company in Action

The experiment wasn’t just theoretical. It involved a live simulation of a company with 13 synthetic employees, real cash flow mechanics, and a public cash countdown. Every decision the AI made was logged and auditable, making the entire process transparent and observable at firmulate.com/live. This setup reveals how AI behaves in real business settings, not just in polished demos.

Why the Gap Matters — and What It Tells Us

The AI that scored highest, gpt-5.6-sol, achieved a perfect 95, thanks to its ability to uncover the buried facts and close the deal. Kimi K3 followed closely at 93, demonstrating the importance of discipline and thoroughness. Meanwhile, Opus 4.8 and Sonnet scored 73 and 88 respectively, showing that even with many learned rules, some discipline slips can lead to missed opportunities.

Implications for Business and AI Adoption

For companies considering AI tools, these findings underline a vital point: performance isn’t just about language fluency or seeming cleverness. It’s about trustworthiness, consistency, and the ability to deliver results under pressure. An AI that refuses manipulation and reads the full context can be more valuable than one that just sounds convincing in a chat.

More broadly, the experiment demonstrates why honest benchmarks matter. They give a transparent, real-world view of AI readiness, helping decision-makers avoid overestimating their tools based on superficial metrics. As firms consider deploying AI in critical functions — from customer management to financial decisions — these benchmarks offer an honest yardstick.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Build vs Buy a Prebuilt AI Workstation

Discover whether building or buying your AI workstation makes sense in 2026. Explore costs, speed, support, and real-world scenarios to make the right call.

Should We Sacrifice Sovereignty For Superior AI Models? Here’s Why

Thorsten Meyer AI says most companies should favor stronger AI models unless law or classified data requires sovereign infrastructure.

Acoustic Dampening, Placement, and the “Rig in the Closet” Setup

Learn how to quiet your workspace with smart placement and acoustic treatment. Discover the secrets to a silent, effective closet studio setup.

AI in Business: Diligence Isn’t Enough to Seal the Deal

A live AI experiment shows that deep effort alone isn’t enough; strategic prioritization and discipline determine whether AI seals the deal or leaves opportunity on the table.