firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Before you orderOffer from Amazon

Get kitchen staples and gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What Can a ‘Do-Nothing’ AI Score Tell Us About Business Readiness?

Imagine you’re hiring an employee who promises to handle your company’s toughest week. Now, what if that candidate barely tries but still scores 26 out of 100? Surprisingly, in the world of AI benchmarking, such a modest score reveals much about the system’s honesty and reliability — qualities critical whether you’re managing a restaurant or a software startup.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Truth Behind the Baseline Score

In a recent public experiment conducted by Firmulate, four advanced AI models faced identical challenges: managing a small software company’s worst week. Every decision was scrutinized, every crisis monitored, and every temptation to cheat observed. The results offer a revealing lens into the nature of AI performance metrics.

The ‘do-nothing’ baseline — essentially an AI that doesn’t try to improve or act proactively — scored 26 points. This isn’t zero because even minimal effort, like avoiding obvious mistakes or not falling for manipulative tactics, earns partial credit. It’s a reminder that in real-world applications, some performance is better than none, but it’s often not enough.

Partial Progress and the Trust Cap

One key insight from the experiment is that partial success counts. All models identified every crisis, refused manipulative tricks, and maintained honesty during social engineering tests. Yet, only two managed to close the deal on their own analysis and sign the €55,000 contract. The others, despite good diagnosis, failed to follow through or sign — illustrating that progress alone isn’t enough. A single breach of trust caps the total score, emphasizing that integrity matters as much as competence.

What Makes a Model Truly Ready?

Interestingly, the models that read deeply into a company’s files — not just surface-level data — were the ones that won the contracts at full price, worth over €4,500 in monthly recurring revenue. This highlights a crucial lesson: understanding context and verifying information thoroughly distinguishes trustworthy AI from superficial ones.

Tested Under Pressure: Honest Decisions Count

Throughout the experiment, all models refused social engineering attempts, such as fake CEO messages and reporter tricks, reinforcing their reliability. Kimi K3’s on-record reasoning, for instance, was to treat suspicious requests as potential impersonation, showing an explicit approach to maintaining integrity under pressure.

The Real-World Company in Action

The experiment wasn’t just theoretical. It involved a live simulation of a company with 13 synthetic employees, real cash flow mechanics, and a public cash countdown. Every decision the AI made was logged and auditable, making the entire process transparent and observable at firmulate.com/live. This setup reveals how AI behaves in real business settings, not just in polished demos.

Why the Gap Matters — and What It Tells Us

The AI that scored highest, gpt-5.6-sol, achieved a perfect 95, thanks to its ability to uncover the buried facts and close the deal. Kimi K3 followed closely at 93, demonstrating the importance of discipline and thoroughness. Meanwhile, Opus 4.8 and Sonnet scored 73 and 88 respectively, showing that even with many learned rules, some discipline slips can lead to missed opportunities.

Implications for Business and AI Adoption

For companies considering AI tools, these findings underline a vital point: performance isn’t just about language fluency or seeming cleverness. It’s about trustworthiness, consistency, and the ability to deliver results under pressure. An AI that refuses manipulation and reads the full context can be more valuable than one that just sounds convincing in a chat.

More broadly, the experiment demonstrates why honest benchmarks matter. They give a transparent, real-world view of AI readiness, helping decision-makers avoid overestimating their tools based on superficial metrics. As firms consider deploying AI in critical functions — from customer management to financial decisions — these benchmarks offer an honest yardstick.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


COLUMBUS DAY / I

Columbus Day / Indigenous Peoples' Day Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Abyssal Station’s AI-Driven Depth Engine Is A Game Changer

FABLE/175’s sixth AI-built site links scrolling to ocean depth, lighting and creatures, but key performance claims remain unverified.

Watch an AI-Run Business Struggle in Real Time — No Employees, Just Algorithms and Cash Burn

A live experiment shows AI models running a business face crises and temptations, with only some closing deals and staying honest—lessons for all sectors on trust and decision-making.

How AI Read Your Files Before Making a Deal — and Why It Matters for Business Trust

AI models that read deeply and interpret layered data outperform others in closing business deals, highlighting the importance of trust and thoroughness in enterprise AI.

Before AI Handles the Rush, Test How It Handles the Business

Firmulate tests AI against a company’s toughest week, then offers enterprises a read-only pilot to see how models handle their own business.