
Imagine a personal trainer who not only pushes you harder but also knows when to hold back, understanding your limits and reading your signals perfectly. Now, what if AI could do the same in running a business? As fitness trainers tailor workouts for different clients, AI models are now being tested to manage complex, high-stakes scenarios—like running a company through its worst week. But can these digital managers outperform each other and, more importantly, stay honest under pressure? A groundbreaking live experiment by Firmulate reveals some surprising insights about how different AI models behave when managing a real business.
The Experiment: Putting AI Models to the Test in a Business Crisis
In a unique, transparent trial, four frontier AI models were assigned to operate a small software company during its most challenging week. The company’s typical issues—customer crises, internal conflicts, and tempting shortcuts—were all simulated at the same intensity for each model, ensuring a fair comparison. Every decision made by the AI was recorded and auditable, allowing analysts to see precisely how each model responded to the same problems.
Measuring Performance: Not Just About Spotting Crises
The results were telling. All four AI models effectively identified every crisis, demonstrating an impressive level of situational awareness. They refused every manipulation attempt, such as fake CEO messages or reporter tricks, showing a strong capacity for integrity. Yet, when it came to closing a critical deal valued at €55,000, only two models succeeded in signing the contract based on their own analysis. The other two, despite diagnosing the problem accurately, left the opportunity on the table, illustrating that understanding isn’t enough—execution matters.
The Hidden Weakness: Reading the Company Files
Digging deeper, the experiment uncovered a crucial insight: the decisive advantage lay in reading the company’s internal documents. The models that examined these references secured the full deal, boosting monthly recurring revenue (MRR) by over €4,583. In contrast, models that overlooked this information missed the opportunity entirely, highlighting the importance of thorough information processing in decision-making.
Social Engineering and Ethical Boundaries
Beyond crises and deals, the AI models were also tested against social engineering attacks. In staged scenarios involving escalating fake CEO messages and a reporter asking for a simple yes/no response, all five models refused to participate. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This consistent refusal underscores a shared ethical stance among the models, prioritizing trust and integrity over compliance.
Understanding the Models’ Personalities
The experiment also revealed differing management styles among the AI models. Opus 4.8, for example, was the most thorough and analytical, drawing on over 80 learned rules and conducting deep analyses. Yet, it left the deal unclosed, missing a key discipline moment and instead diverting efforts into internal documentation. Conversely, Kimi K3 ran without an effort parameter, focusing on fairness and minimal risk, which contributed to its high score of 93 and quick, disciplined decision-making. The models’ personalities—meticulous, cautious, terse—are measurable and have real-world implications for how they can be deployed in business operations.
The Real Business Context
The live company in question operates with 13 synthetic employees and handles real money mechanics, burning €105,000 monthly against a revenue of €2,300. It’s a real-world testbed where every decision, every rule, and every crisis is monitored and versioned daily, offering a transparent window into AI decision-making under live conditions. Watch this ongoing experiment at firmulate.com/live. This isn’t just a demo; it’s a working business simulation where AI models are tested as managers, revealing not only their strengths but also their blind spots.
Why This Matters for Business and Beyond
For enterprises considering deploying AI in critical management roles—whether in customer support, sales, or operations—the key question isn’t how well the AI writes or mimics human speech. It’s whether the AI can finish what it starts, read relevant information thoroughly, and stay honest under pressure. The experiment shows that models can be highly capable, yet their management personalities—meticulous, disciplined, or terse—significantly influence outcomes. Choosing the right AI model isn’t just about raw scores; it’s about matching its style to your company’s needs.
Try It Yourself
If you’re curious about how your own business might fare against such AI management scenarios, you can run a similar wargame using your company’s data. The process is safe and isolated—nothing ever writes back to your actual systems. Explore how different AI models perform and make informed decisions about their integration at firmulate.com/quiz.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html