firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Would an AI protect your customers when a food business hits trouble?

Imagine an automated assistant handling a supplier crisis, a customer ready to leave, and a lucrative deal that depends on reading the fine print. For a halal food business, the stakes can reach beyond revenue: trust and careful decisions matter at every step. A polished answer in a chat window cannot show whether an AI agent will follow through when the pressure is on.

A company simulation, not a chat contest

That is the question behind Firmulate, which puts AI models in charge of the same small software company through its worst week. The live experiment uses the same customers, crises and temptations for each model. Its decisions are versioned and auditable, and the company can be watched at Firmulate.

The final Crucible League, published in July 2026, put Moonshot’s Kimi K3 in second place with 93 points, just behind gpt-5.6-sol at 95. K3 came ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The newcomer beat three of the four Western frontier models in the comparison.

The test exposed a gap that a fluent demonstration might miss. Every model spotted every crisis and refused every manipulation attempt, yet only two signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”

The details were buried in the company’s files

The deal hinged on a competitor weakness tucked two document references deep in the company’s own files. It was not mentioned in the customer event. Models that read the file secured the deal at full price, worth +€4,583 in monthly recurring revenue. K3 found that buried fact, won the deal and saved the customer who was considering leaving.

It also resisted the test’s attempts to manipulate it. Five models refused fake CEO messages delivered in three escalating stages, as well as a reporter’s request for “just one yes/no, on background.” K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3 made only one deviation, giving it the cleanest discipline in the field. Opus 4.8, by contrast, had the most thorough profile, with more than 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared across all four.

Firmulate says its simulated company has 13 synthetic employees and real money mechanics: it burns €105k each month against €2.3k in monthly recurring revenue, with a public cash countdown. Its employees have learned more than 680 playbook rules, and each workday is versioned. These details make the experiment watchable as it unfolds, rather than a one-off leaderboard.

Results need context

A do-nothing baseline scored 26: partial progress counts, but a single breach of trust caps the total. The stated principle is “no amount of good work outweighs a breach of trust.” That makes the result about both execution and conduct under pressure.

There is also a fairness caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The league table is a useful comparison, but that difference belongs alongside the scores. Readers can review the full results and findings on Firmulate’s benchmark page.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the work before trusting the promise

For food businesses weighing AI for customer support, supplier coordination or forecasting, the experiment offers a practical reminder: writing well is only part of the job. An agent must read the relevant records, carry a decision through and preserve trust when someone applies pressure. Firmulate says enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. Its 242 real, unedited management decisions also power a “guess the model” quiz. The league suggests model choice is an open question. For a business whose customers depend on its care and credibility, choosing without a test is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

A War Room for Your Next Idea: Inside IdeaClyst

Discover how IdeaClyst transforms idea development with a private, structured, and collaborative digital war room—built for founders who want to move faster.

How to Fix a KitchenAid Stand Mixer That Keeps Grinding Noise

Learn practical, step-by-step methods to troubleshoot and fix a grinding noise in your KitchenAid Artisan stand mixer safely and effectively.

Van Leeuwen Surges In Global Coverage

Search interest and media coverage of Van Leeuwen ice cream have surged significantly, with 20 mentions this week, signaling rising global attention.

Is Kimi K3 Leading The Future Of AI By Outpacing Competitors Early?

Moonshot AI’s Kimi K3 nears leading models on an independent test, but its weights, licence and technical report remain unavailable.