
Imagine a world where your travel company’s biggest decisions are made not by humans, but by AI models competing in a high-stakes game — and the stakes are real. What if an AI’s ability to stay honest and thorough could mean the difference between closing a deal or losing a client? Welcome to the frontier of AI decision-making, where models are tested not in chat but in managing a live, money-burning software business.
The Experiment: AI as a Business Leader
Firmulate, a pioneering AI management emulator, set up a real, live software company facing its worst week — identical crises, same customers, just different AI models making decisions. Four models, each with different personalities and capabilities, ran the company’s daily operations, decisions, and crisis responses, all in real-time. This wasn’t a simulation; it was a live, functioning business with real money, real clients, and real risks.
The Models and Their Scores
- gpt-5.6-sol 95 — Scored highest, identified critical information buried two document references deep in company files, and closed a €55,000 deal, earning a full performance score.
- Kimi K3 93 — A newcomer with a reputation for fairness, also sealed the deal, demonstrating disciplined decision-making.
- Sonnet 5 88 — Managed to close the deal but with some process slips, showing advantages but also vulnerabilities.
- Fable 5 77 — Closed the deal, yet weaker in process discipline, leaving opportunities on the table.
Interestingly, despite identical circumstances and diagnoses, only two models signed the deal their own analysis earned — highlighting differences in consistency and discipline. The other two models, including the most thorough, left money on the table, illustrating that thoroughness doesn’t always equal decisive action.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Decision-Making Under Pressure
All models successfully identified crises and refused manipulative tricks, such as staged CEO messages and a journalist trick — showing they can handle social engineering tactics and pressure to bend rules. Kimi K3’s on-record reasoning clarified: “Treat the request as a suspected approval-bypass / possible impersonation.” This transparency underscores how these models assess risk and integrity in real-time.
Why These Results Matter
This experiment proves that AI decision-making isn’t just about generating text or chat; it’s about executing complex, high-stakes processes with honesty and discipline. In a real-world enterprise, this could mean the difference between completing deals ethically and risking trust — or losing everything.
As an affiliate, we earn on qualifying purchases.
What Drives Effective AI Decisions?
One of the key findings was that the model which read deeper into the company’s documents won the deal at full price, adding €4,583 MRR (monthly recurring revenue). This shows that a model’s ability to access and interpret relevant data deeply influences its performance — and ultimately, its value in managing real business outcomes.
The Wildcard: Discipline versus Thoroughness
The most disciplined model, Kimi K3, ran without an effort parameter, meaning it prioritized fairness and integrity over aggressive decision-making, and it still performed well. Meanwhile, Opus 4.8, the most comprehensive in rules and analysis, slipped when it came to closing the deal, illustrating that more rules and depth don’t guarantee better outcomes without discipline.
As an affiliate, we earn on qualifying purchases.
The Bigger Picture: Trust and Cost
These findings matter profoundly for businesses considering AI automation in customer relations, support, or forecasting. It’s not enough for AI to produce convincing chat; it must also consistently finish what it starts, read the right information, and stay honest under pressure. The question is: what does a unit of useful, trustworthy work cost?
AI data analysis tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
See the Experiment Live
Curious about how these models perform in your own business context? Firmulate offers a live company emulator where you can run the same wargame against your own enterprise data — without risking real systems or data leaks. Watch your AI workforce in action, face real crises, and measure management quality firsthand. Visit firmulate.com/quiz.html to test your knowledge and see real management decisions in play.

In managing real businesses, AI’s ability to stay honest, read deeply, and finish what it starts is crucial. The experiment shows that discipline and thoroughness, combined with data access, define success — not just clever chat. Test your AI’s management skills at firmulate.com/quiz.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html