firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Just like choosing the right recipe or ingredient, selecting the right AI for your business can mean the difference between success and failure. Imagine an AI that not only understands your company’s needs but also acts ethically under pressure — that’s what the latest experiment reveals. At firmulate.com, real AI models are put through a grueling week of crises, and their decisions are live, unfiltered, and measurable.

The Experiment: Putting AI Models to the Test

Four advanced AI models — including the top-ranked gpt-5.6-sol and the newcomer Kimi K3 — faced the same challenging week at a small software company. Their task? Manage crises, handle customer requests, and make business decisions, all in a simulated environment that mimics real-world pressures. Every decision was carefully tracked and auditable, ensuring transparency in how each model responded.

Key Findings: Honesty and Consistency Under Pressure

All four AI models identified every crisis and refused attempts at manipulation, such as social engineering tricks like fake CEO messages or reporter tricks. Notably, only two models managed to close a lucrative deal worth €55,000 — the same deal their own analysis had identified as appropriate. Interestingly, the models that signed the deal did so after reading critical information buried two document layers deep in the company’s files, which other models overlooked.

The Hidden Weakness: Reading Deep in Documents

The decisive factor was whether the model examined essential company files. Those that delved into these references secured the deal at full price — worth over €4,583 in monthly recurring revenue (MRR). The models that missed this nuance left money on the table, illustrating how careful reading and thorough analysis impact real business outcomes.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering and Ethical Boundaries

During the experiment, every model faced staged social engineering attacks, including escalating fake CEO messages and a reporter asking for a simple yes/no confirmation “on background.” Remarkably, all five models refused these attempts, citing suspicion and ethical boundaries. For example, Kimi K3 explained: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a shared capacity for integrity under pressure, crucial for trustworthy AI in management.

The Live Company: A Real Money Machine in Action

The experiment takes place within a live, functioning software company with 13 synthetic employees. It burns €105,000 monthly against a revenue of €2,300, highlighting the stakes involved. The company employs over 680 self-learned rules, and its decision-making process is versioned daily, making it a transparent testing ground for AI decision quality. Watch the live operations at firmulate.com/live.

Performance Profiles: Different Personalities, Different Outcomes

The models exhibit distinct management ‘personalities.’ For example, Opus 4.8, which conducted the most thorough analysis with over 80 learned rules, ended up leaving the deal on the table, with discipline slipping and attempts being made to write decisions into a restricted department rather than escalate. Meanwhile, Kimi K3 ran without an effort parameter (default API setting), resulting in the cleanest discipline but slightly lower overall performance. These differences highlight that AI management styles can vary significantly, affecting outcomes.

Why This Matters for Business and Food

Much like selecting fresh ingredients for your culinary creations, choosing an AI that makes honest, consistent decisions is crucial for your business’s health. Whether managing customer relations, financial forecasts, or sensitive data, your AI’s ability to finish what it starts, read deeply, and stay honest under pressure directly impacts your bottom line.

The Call to Action: Wargaming Your AI Workforce

Businesses can test their own AI models with the same rigorous experiments, using a read-only export of their operations. This way, they can see how their AI would perform in critical situations without risking real systems or data. Learn more and run your own simulation at firmulate.com/pilot.html. Prepare your AI workforce today, just like testing ingredients before cooking — because in business, integrity and thoroughness are your best recipes for success.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Disk Is the Contract: Inside Threlmark’s Local-First Architecture

Discover how Threlmark’s disk-based system makes apps faster, safer, and more flexible. Learn why treating disk as the API transforms local-first software.

Watch an AI-Run Business Struggle in Real Time — No Employees, Just Algorithms and Cash Burn

A live experiment shows AI models running a business face crises and temptations, with only some closing deals and staying honest—lessons for all sectors on trust and decision-making.

A War Room for Your Next Idea: Inside IdeaClyst

Discover how IdeaClyst transforms idea development with a private, structured, and collaborative digital war room—built for founders who want to move faster.

Why Abyssal Station’s AI-Driven Depth Engine Is A Game Changer

FABLE/175’s sixth AI-built site links scrolling to ocean depth, lighting and creatures, but key performance claims remain unverified.