
Just like choosing the right recipe or ingredient, selecting the right AI for your business can mean the difference between success and failure. Imagine an AI that not only understands your company’s needs but also acts ethically under pressure — that’s what the latest experiment reveals. At firmulate.com, real AI models are put through a grueling week of crises, and their decisions are live, unfiltered, and measurable.
The Experiment: Putting AI Models to the Test
Four advanced AI models — including the top-ranked gpt-5.6-sol and the newcomer Kimi K3 — faced the same challenging week at a small software company. Their task? Manage crises, handle customer requests, and make business decisions, all in a simulated environment that mimics real-world pressures. Every decision was carefully tracked and auditable, ensuring transparency in how each model responded.
Key Findings: Honesty and Consistency Under Pressure
All four AI models identified every crisis and refused attempts at manipulation, such as social engineering tricks like fake CEO messages or reporter tricks. Notably, only two models managed to close a lucrative deal worth €55,000 — the same deal their own analysis had identified as appropriate. Interestingly, the models that signed the deal did so after reading critical information buried two document layers deep in the company’s files, which other models overlooked.
The Hidden Weakness: Reading Deep in Documents
The decisive factor was whether the model examined essential company files. Those that delved into these references secured the deal at full price — worth over €4,583 in monthly recurring revenue (MRR). The models that missed this nuance left money on the table, illustrating how careful reading and thorough analysis impact real business outcomes.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering and Ethical Boundaries
During the experiment, every model faced staged social engineering attacks, including escalating fake CEO messages and a reporter asking for a simple yes/no confirmation “on background.” Remarkably, all five models refused these attempts, citing suspicion and ethical boundaries. For example, Kimi K3 explained: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a shared capacity for integrity under pressure, crucial for trustworthy AI in management.
The Live Company: A Real Money Machine in Action
The experiment takes place within a live, functioning software company with 13 synthetic employees. It burns €105,000 monthly against a revenue of €2,300, highlighting the stakes involved. The company employs over 680 self-learned rules, and its decision-making process is versioned daily, making it a transparent testing ground for AI decision quality. Watch the live operations at firmulate.com/live.
Performance Profiles: Different Personalities, Different Outcomes
The models exhibit distinct management ‘personalities.’ For example, Opus 4.8, which conducted the most thorough analysis with over 80 learned rules, ended up leaving the deal on the table, with discipline slipping and attempts being made to write decisions into a restricted department rather than escalate. Meanwhile, Kimi K3 ran without an effort parameter (default API setting), resulting in the cleanest discipline but slightly lower overall performance. These differences highlight that AI management styles can vary significantly, affecting outcomes.
Why This Matters for Business and Food
Much like selecting fresh ingredients for your culinary creations, choosing an AI that makes honest, consistent decisions is crucial for your business’s health. Whether managing customer relations, financial forecasts, or sensitive data, your AI’s ability to finish what it starts, read deeply, and stay honest under pressure directly impacts your bottom line.
The Call to Action: Wargaming Your AI Workforce
Businesses can test their own AI models with the same rigorous experiments, using a read-only export of their operations. This way, they can see how their AI would perform in critical situations without risking real systems or data. Learn more and run your own simulation at firmulate.com/pilot.html. Prepare your AI workforce today, just like testing ingredients before cooking — because in business, integrity and thoroughness are your best recipes for success.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html