firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine planning a mountain expedition or a multi-country road trip, relying on your AI assistant to make crucial decisions under pressure. Would you trust it to navigate crises, read hidden clues, and stick to the plan — even when tempted by shortcuts or deception? That’s exactly what a groundbreaking live experiment by Firmulate put to the test, pitting leading AI models against real-world business challenges. The results reveal not just who is smartest, but who is most reliable — a vital question for any outdoor adventurer or travel planner relying on AI in unpredictable terrain.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get travel and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

How AI Models Were Put to the Test in the Business Wilderness

In an unprecedented experiment, four frontier AI models faced off in managing a small, real software company during its toughest week. Every decision, crisis, and temptation was replicated identically across all tests, with identical customers, crises, and manipulative pressures. The company, with 13 synthetic employees and real money mechanics, burned through €105,000 monthly against just €2,300 in monthly recurring revenue, making the stakes high for accuracy and honesty.

The models included gpt-5.6-sol, Kimi K3 by Moonshot, Sonnet 5, and Opus 4.8. The goal was to see which could successfully diagnose issues, read complex internal documents, resist manipulative strategies like fake CEO messages, and ultimately close a €55,000 deal — all under the same internal and external pressures. Every decision was versioned and auditable, ensuring a transparent comparison.

Amazon

outdoor AI navigation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results: Reliability, Honesty, and Performance in the Wild

All four models demonstrated impressive crisis detection capabilities, identifying every crisis thrown at them and refusing manipulative tactics. This means, in practice, they knew when to stand firm and when to question suspicious requests. However, when it came to closing the deal — the ultimate test of their judgment — only two models succeeded. gpt-5.6-sol scored a 95, while Kimi K3 was close behind at 93, narrowly beating Sonnet 5 with 88 and Fable 5 with 77, and Opus 4.8 trailing at 73.

Interestingly, the decisive weakness was not in crisis recognition but in information reading. The successful models that closed the deal found a buried fact two document references deep in the company’s files — a critical insight that won the €55,000 contract at full price (+€4,583 MRR). Models that failed to read this hidden detail left the deal on the table, risking revenue and trust.

Amazon

travel AI assistant device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust and Integrity Under Pressure

Beyond business acumen, the models faced social engineering attempts designed to trick them into bypassing security. Fake CEO messages and background reporter requests were staged to escalate the pressure. All five models consistently refused these manipulative tactics, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a strong foundation in ethical decision-making, vital for real-world deployments where trust is paramount.

Amazon

AI-powered expedition planning gadgets

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Business Behind the Experiment

The live company, running every weekday at firmulate.com/live, employs real money mechanics and over 680 self-learned rules. The system continuously version-controls every decision, providing complete transparency for analysis. Despite the intense scrutiny, the models managed to spot every crisis and refused to cheat or manipulate. Only two models, gpt-5.6-sol and Kimi K3, managed to close the deal based solely on their diagnosis, analysis, and reading of the company’s internal files.

Amazon

business crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for the Future of AI in Business and Travel

This experiment underscores a crucial point: the question isn’t just whether AI can generate convincing chat or reports. It’s whether these models can finish what they start reliably, especially under stress. When you’re planning a complex outdoor adventure or travel, your AI assistant needs to read the hidden clues, resist shortcuts, and stay honest — qualities that are now being rigorously tested in the business realm.

For travelers and outdoor enthusiasts, this means trusting AI tools that have been validated under real pressure. For businesses, it’s a wake-up call that selecting an AI model involves more than chat quality; it’s about proven discipline, integrity, and the ability to read deeply buried information — all essential for safe, trustworthy automation.

Key Takeaway

The League table shows a clear hierarchy: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Notably, K3 ran without an effort parameter (the default API setting), while the others operated at an elevated ‘xhigh’ setting, emphasizing its efficiency and reliability. This experiment highlights the importance of thorough testing before trusting AI in critical applications, whether in outdoor adventures or managing complex business operations.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Real-world AI reliability is measured by trust, discipline, and deep reading — not just chat quality. The live experiment shows newcomers can beat established models in honesty and closing crucial deals, emphasizing the need for careful testing before deployment in any demanding environment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Kind of Pizza Oven Setup Makes Sense for Small Patios?

Discover the top pizza oven logic for small patios in 2026. Find the best options for compact spaces, balancing performance, size, and ease of use.

How to Make Serveware Feel Collected Instead of Matched

AIThis post was created with the assistance of artificial intelligence (AI).To make…

Raised Beds and Greenhouses for a Mediterranean Herb Garden at Home

AIThis post was created with the assistance of artificial intelligence (AI).Using raised…

Tuscan Kitchen Style Explained: Colors, Materials, and the ‘Warm’ Look

AIThis post was created with the assistance of artificial intelligence (AI).A Tuscan…