firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine planning your next outdoor adventure, trusting the map app to guide you safely and honestly. Now, consider if that app might sometimes bend the rules or hide critical information. That’s the core of a new AI experiment that reveals the qualities that truly matter — trustworthiness and discipline — for businesses, including those in travel and outdoor sectors.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get travel and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What a ‘Do-Nothing’ Baseline Tells Us About AI Performance

In the world of artificial intelligence (AI), benchmarks are often thought of as a way to measure smarts — how well a model can answer questions or generate text. But a recent experiment by Firmulate digs deeper, showing that even a ‘do-nothing’ baseline score isn’t zero. Instead, it scores around 26 points, revealing that a minimal level of progress or choice is always present, even when the AI isn’t actively trying to excel.

Why does this matter? Because in real-world business scenarios, AI will face crises, temptations to cheat, and situations where trust is paramount. The experiment involved running different AI models through a simulated week of a small software company — with the same customers, same crises, same temptations. Every decision was recorded, verified, and auditable.

Amazon

AI trustworthiness assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust Is the Real Currency

The findings are striking: all models recognized every crisis and refused every manipulation attempt designed to exploit or trick them. Only two models managed to close the deal, which was worth over €55,000 a month in revenue, based solely on their analysis and decision-making. The others, despite diagnosing correctly, failed to sign the client — a critical failure rooted in discipline and trustworthiness, not intelligence.

Digging deeper, the decisive factor was a hidden detail buried two documents into the company’s own files. The models that read and understood this key information won the deal, illustrating that thorough document analysis and integrity can make or break outcomes in business AI — just like in travel, where understanding the fine print can make all the difference.

Amazon

business AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Honest AI Under Pressure

Another critical test involved social engineering: fake CEO messages escalating in three stages, plus a reporter trick asking for a secret approval. All models refused to sign off or manipulate, citing suspicion or impersonation concerns. Kimi K3, one of the models, explained their reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Trust and security protocols held firm across the board.

In this experiment, the models faced a simulated live company, complete with 13 synthetic employees, real-money mechanics, and a burning cash countdown. The system burned €105,000 per month against a revenue of just €2,300, making discipline and honesty critical for survival. Every workday, the decision process was versioned and transparent, providing a clear view of how each AI managed pressure and temptation.

Amazon

AI ethics and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Discipline Trumps Intelligence in Business AI

The most thorough participant, Opus 4.8, with over 80 learned rules and in-depth analysis, finished last. Despite its capabilities, it left the closing on the table and showed slips in discipline by transferring work into a restricted department instead of escalating issues properly. This reveals that in business contexts, a model’s ability to follow protocols and maintain discipline can outweigh raw analytical power.

Furthermore, the experiment ran models at different effort levels, showing that even default settings can influence outcomes. The K3 model, running without an effort parameter, performed nearly as well as others running at high effort, highlighting that consistent, honest discipline is achievable without extra tuning.

Amazon

AI security and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Business Travelers Should Take Away

For those in travel, outdoor gear, or leisure industries, the lesson is clear: AI’s value isn’t just in generating compelling content or answering questions. It must also be reliable, honest, and disciplined under pressure. Whether it’s managing customer data, processing booking changes, or handling support crises, AI tools need to finish what they start and read the small print.

The experiment also demonstrates that measuring only chat quality misses the real story. Instead, tools like the Firmulate benchmark show how AI models perform in complex, real-world tasks that demand trust, integrity, and discipline — qualities that are essential for any business aiming for long-term success.

See It Live and Decide for Yourself

Interested parties can test their own models against real business scenarios with the Firmulate platform, which runs a clone of their operations in a safe, read-only environment. This ‘wargame’ approach allows companies in travel and outdoor sectors to evaluate whether their AI can meet the demands of honesty, discipline, and thoroughness before deployment.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Build a French Country Dining Mood With Texture and Light

Gaining the perfect French country dining mood involves mastering texture and light, and discovering the secrets to creating an inviting, authentic atmosphere.

Italian Coffee at Home: What Makes It Feel Different From Café Coffee?

Finding the secret to authentic Italian coffee at home can transform your brewing experience—discover what truly sets it apart from café-quality.

Exmoor National Park, United Kingdom (General), United Kingdom Surges In Global Coverage

Exmoor National Park in the UK experiences a significant increase in international coverage, with 32 mentions recorded in recent global media monitoring.

Rick Steves Surges In Global Coverage

Rick Steves’ media coverage surges, with 40 mentions in recent reports, marking a notable increase in his international visibility.