firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Imagine a travel guide that meticulously plans every step, yet still leaves the journey incomplete. In the world of AI-driven business decisions, thoroughness alone doesn’t guarantee success. Recent experiments reveal that even the most diligent AI models can stumble at the final hurdle, emphasizing that prioritization and focus often outweigh sheer volume of effort.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Experiment: Simulating a Tough Week for a Software Company

At the heart of this investigation is a live, real-world simulation conducted by Firmulate, an innovative platform that models entire companies as AI agents. In this experiment, four advanced AI models were tasked with navigating a challenging week for a small software firm—facing the same customers, crises, and temptations to cut corners. Every decision was carefully versioned and auditable, providing a transparent view of their choices and behaviors.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Models and Their Performance

Among these was Opus 4.8, a profile renowned for its meticulous approach. It learned over 80 rules and conducted the deepest analyses among its peers. Yet, despite its thoroughness, Opus finished last in the league table, with a score of 73 out of 100. The top performer, gpt-5.6-sol, scored a perfect 95, while Kimi K3 and Sonnet 5 followed closely with 93 and 88 respectively.

All four models identified every crisis and refused every manipulation attempt, such as social engineering tactics like fake CEO messages or subtle requests to bypass approval processes. This indicates that they were equally aware of what was ethical and what was not. However, only two models managed to close the deal worth €55,000—a full-price contract that their own analyses had justified. The other two, including Opus, failed to sign despite having the correct diagnosis and pitch.

The Hidden Weakness: Reading Deeper in the Files

The critical difference was not in surface-level decision-making but in how deeply each model read into the company’s own documents. The models that succeeded in closing the deal accessed information buried two document references deep within the company’s files—details that were pivotal for making the final decision. Conversely, Opus 4.8, despite its thorough analysis, did not retrieve this crucial information, leading to a missed opportunity that was worth an additional €4,583 in Monthly Recurring Revenue (MRR).

Amazon

business AI analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Insights for Business and AI Deployment

This experiment underscores a vital lesson: diligence alone does not guarantee impact. An AI model can be exhaustive in its analysis but still fail at the critical moment if it does not prioritize key information. The models that succeeded demonstrated a capacity for targeted focus—reading just enough to make the right decision—rather than attempting to process every available detail indiscriminately.

Security and Trust Under Pressure

In the social engineering component of the test, all models refused to be manipulated through staged fake messages and reporter tricks. For example, Kimi K3 explicitly treated suspicious requests as potential impersonation attempts, reflecting a cautious, security-first mindset. This consistency in refusal highlights that current models are capable of maintaining integrity even under pressure.

Amazon

AI document analysis platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Company: A Real-World Experiment

Beyond the simulation, Firmulate runs a live, observable company emulation with 13 synthetic employees managing real money mechanics—burning €105,000 monthly against a revenue of €2,300. Every operational rule is self-learned and versioned daily, creating a transparent environment where decision-making processes are scrutinized and improved continuously. Viewers can observe the company’s performance at firmulate.com/live.

Amazon

AI security and trust solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Broader Takeaway: Focus Over Volume in AI Decisions

The core lesson from the experiment is clear: in both AI and business, discipline and thoroughness are not enough if they are not paired with sharp prioritization. The most successful models read selectively, focus on critical facts, and avoid getting bogged down by details that do not influence the outcome.

This insight is especially vital as AI begins to touch areas like customer support, CRM, and forecasting—domains where quick, accurate, and honest decisions directly impact profitability and trust. The question for enterprises is no longer whether their AI can produce polished responses, but whether it can finish what it starts, stay honest under pressure, and extract the most valuable insights efficiently.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI in Business: Can Machines Lead Under Pressure or Just Score in Tests?

Leading AI models can identify crises and refuse manipulation, but only those reading deeply and staying disciplined win crucial deals. Real-world resilience matters most.

PlayStation Can Delete All Your Digital Games After 3 Years Of Inactivity (EU)

Sony PlayStation has confirmed it will delete all digital games from accounts inactive for over three years in the European Union, effective immediately.

Seagate Technology Surges In Global Coverage

Seagate Technology experiences a surge in worldwide media coverage, with 25 mentions reported in recent monitoring data, highlighting increased public and industry interest.

Microsoft Surges In Global Coverage

Microsoft experiences a surge in worldwide media mentions, with GDELT reporting 135 mentions in recent coverage, indicating increased global attention.