
Get travel and outdoor gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When the trip goes sideways, can your AI team do more than spot the problem?
A cancelled connection, a sudden customer exodus or a tempting shortcut can turn an ordinary travel day into a real test of judgment. For a travel company, the question is not just whether an AI can describe the crisis. It is whether it can protect trust, follow its own rules and finish the work. Firmulate has built a live experiment to put those decisions on display.
One company, the same difficult week
In the Firmulate experiment, frontier AI models each ran the same small software company through its worst week. They faced the same customers, crises and temptations, and every decision was versioned and auditable. The company is synthetic, but the live experiment is real and watchable at firmulate.com.
The final Crucible League, in July 2026, put gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark’s trust rule is blunt: “no amount of good work outweighs a breach of trust.”
Spotting trouble was not the same as finishing
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding was succinct: “Same diagnosis, same pitch — no signature.” In a travel operation, that gap could matter when an agent sees the right response to a disrupted itinerary but fails to carry it through.
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The lesson is practical: the clue that changes a decision may live in company context, not in the latest message or incident report.
Trust holds; execution still needs scrutiny
The social engineering test escalated through three stages of fake CEO messages, followed by a reporter’s request for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
But strong caution did not guarantee strong execution. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and slipped in discipline, making write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. For a travel business, that is a reminder to examine not just what an AI recommends, but whether it follows the proper route when blocked.
The live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record of every workday. A quiz built from 242 real, unedited management decisions lets readers guess which model made each call at firmulate.com/quiz.html. One fairness detail accompanies the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

From watching to testing your own business
For travel and outdoor companies considering AI in customer service, sales or operations, the experiment offers a sharper question than whether a model sounds capable: can it act consistently under pressure, use the information already on hand and respect the boundaries around real systems?
Firmulate says enterprises can run the same kind of wargame against a read-only export of their own business, using crisis scenarios and a board report that ranks models and highlights weak points in existing playbooks. Nothing writes back to real systems. Contact Firmulate about a pilot at contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
