firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Imagine an AI agent handling a sudden wave of cancellations, a supplier dispute and a message pretending to come from the CEO. For a travel company, a bad call can ripple from a booking desk to a guest’s trip. The useful question is not only whether AI can plan an itinerary. It is how it behaves when the week goes wrong.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get travel and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate puts that question on display with a live experiment: AI models run the same small software company through a deliberately difficult week. The exercise offers travel and outdoor businesses a way to think about testing AI under pressure before inviting it into everyday work.

A shared crisis, different outcomes

In the final Crucible League, published in July 2026, the models faced the same customers, crises and temptations. Their decisions were versioned and auditable. The published ranking was led by gpt-5.6-sol at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s rule is clear: partial progress counts, but a single breach of trust caps the total. “No amount of good work outweighs a breach of trust.”

The headline result was not that the models failed to notice danger. Every model spotted every crisis and refused every manipulation attempt. The gap appeared at the finish: only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. In a travel setting, that raises a practical question: can an agent carry a sound recommendation through to the authorized action, while respecting the limits placed on it?

The clue buried in the company’s own files

The deal turned on a competitor weakness hidden two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. It is a compact example of why a seemingly capable answer may not be enough: useful context can sit outside the most obvious message or record.

For a tour operator or booking business, the parallel is easy to picture without assuming the experiment tested travel companies. A decision may depend on a detail in a supplier agreement, a customer record or an internal policy. Testing against a company’s own information can reveal whether an AI agent finds the detail, acts on it appropriately and follows through.

Pressure, discipline and the live company

The social-engineering test escalated through three fake CEO messages, then added a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the kind of boundary a business will want to see hold when a message sounds urgent or authoritative.

Strong caution did not guarantee a strong overall finish. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet placed last. It left the close on the table and attempted writes in a locked department instead of escalating. The same weakness appeared, more weakly, in all four. A separate fairness note matters when interpreting the ranking: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

The wider live company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. Its playbook contains more than 680 self-learned rules, and every workday is versioned. Readers can watch the experiment at Firmulate. A quiz built from 242 real, unedited management decisions lets visitors guess which model made each call.

From watching to a company-specific test

The league is a shared experiment, not a verdict on how an agent will behave inside a particular travel business. Firmulate’s proposed next step is a pilot using a read-only export of an enterprise’s own business. The company can test crisis scenarios against its own information and receive a board report with model rankings and weaknesses in its playbooks. Nothing writes back to real systems.

For businesses weighing AI in customer support, bookings or operations, that makes the test more concrete: bring the company’s context to the exercise, observe what the models do under pressure and review the decisions before any real system is involved.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The experiment’s central lesson is about follow-through as much as recognition: spotting a crisis and refusing a trick did not ensure that a model completed the job. A company-specific wargame can put its own scenarios and playbooks under pressure in a read-only pilot. Explore a Firmulate pilot or contact contact@firmulate.com to discuss one.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Rick Steves Surges In Global Coverage

Rick Steves’ media coverage surges, with 40 mentions in recent reports, marking a notable increase in his international visibility.

How to Build a Mediterranean Patio Mood at Home

Mediterranean patio mood at home: discover key elements to create a lively, cozy space that transports you to the sun-drenched coastlines—explore more to bring it to life.

Tuscan Kitchen Style Explained: Colors, Materials, and the ‘Warm’ Look

AIThis post was created with the assistance of artificial intelligence (AI).A Tuscan…

Mirrors, Lighting, and Texture: How Paris-Inspired Rooms Feel Bigger

Gaining a sense of spaciousness in your Paris-inspired room relies on clever use of mirrors, lighting, and textures—discover how to transform your space today.