
For travelers and outdoor enthusiasts, the promise of AI is often seen in smart navigation or personalized recommendations. But what if AI could actually run a business — making critical decisions under pressure? Recent experiments reveal that the latest frontier models are not just chatty assistants but capable managers, with one newcomer outperforming seasoned AI rivals in a real-world simulation.
Get travel and outdoor gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Test: A Business Week in the Trenches
To gauge their true management skills, four advanced AI models faced the same challenge: run a small software company through its worst week. This test, known as the Crucible League, was no superficial demo. It involved identical customers, crises, and temptations, with every decision documented and auditable. The goal? See which AI could spot problems, resist manipulation, and close a crucial €55,000 deal.
As an affiliate, we earn on qualifying purchases.
Results That Make Waves
The scores tell the story: gpt-5.6-sol led with a perfect score of 95, followed closely by the Moonshot newcomer Kimi K3 at 93. The established models—Sonnet 5, Fable 5, and Opus 4.8—lagged behind, with scores of 88, 77, and 73 respectively. Notably, the do-nothing baseline scored a mere 26, emphasizing the level of progress made by these AI players.
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Gave K3 the Edge?
While all models recognized crises and refused manipulative tricks like fake CEO messages, Kimi K3 distinguished itself by uncovering a buried piece of critical information deep within the company’s internal files—something the others missed. This led to the successful closing of the deal at full price, generating an additional €4,583 in Monthly Recurring Revenue (MRR).
AI cybersecurity and fraud detection solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Honest Under Pressure
In scenarios designed to test their integrity, all four models refused to be duped by staged social engineering attempts, including escalating fake CEO messages and a reporter trick. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline in honesty under stress underscores a key advantage in AI management agents.
As an affiliate, we earn on qualifying purchases.
The Significance of the Findings
The experiment underscores a vital point: in real business operations, the ability to read internal files, resist manipulation, and make disciplined decisions is more valuable than just natural language prowess. The models ran the same small company, faced the same crises, and yet the newcomer K3 delivered the best performance, beating established giants in a field that many assumed was still out of reach for AI.
About the Live Experiment
Put into practice, this AI-driven management is not just a lab experiment. The live company, with 13 synthetic employees and real financial mechanics, burns €105,000 monthly against a modest €2,300 MRR. Every decision made by the models is recorded, versioned, and observable at firmulate.com/live. This ongoing testing ground demonstrates that AI models can manage a real business under real pressures, providing a transparent measure of their management qualities.
The Race Continues
While Opus 4.8, with its in-depth analysis and over 80 learned rules, finished at the bottom—failing to close the deal and slipping into process slips—the gap between the newcomers and veterans is narrowing. The league table shows a clear advantage for models that prioritize integrity and thoroughness without sacrificing decisiveness.
Fairness and Transparency
It’s worth noting that Kimi K3 ran without an effort parameter (the API default), while the others ran at xhigh. This ensures a fair comparison, emphasizing that the superior performance is rooted in the model’s inherent capabilities, not just configuration choices.
Broader Implications
This experiment is a wake-up call for enterprises contemplating AI integration. The key question is not just whether an AI writes well, but whether it can finish what it starts, read internal files, resist manipulation, and deliver measurable work. The league table is open, and choosing a model without your own test might be a gamble.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
