firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Why Trust Matters More Than Scores in AI Performance

Just like in your workout routine, where consistency and discipline outweigh flashy moves, evaluating AI systems requires more than quick wins or high scores. A recent experiment by Firmulate sheds light on what really separates a reliable AI from one that’s just good at the surface.

The Firmulate Benchmark: Testing AI Under Real-World Pressure

Imagine putting AI models through a week of managing a small software company facing genuine crises—customers demanding attention, financial pressures, and ethical dilemmas. That’s exactly what the Firmulate experiment did, using a live, watchable platform where AI models navigated the same tough decisions, with every move recorded and auditable.

The goal? To measure management quality—how well these AI systems handle crises, stay honest, and deliver results—not just their ability to generate convincing chat responses. The models faced the same scenarios, the same customer demands, and the same temptations to cheat or cut corners.

What the Scores Say About Trust and Discipline

  • The top-performing AI, gpt-5.6-sol, scored an impressive 95 points, successfully identifying hidden information that sealed a crucial deal.
  • Close behind was Kimi K3, scoring 93, which demonstrated the cleanest discipline, refusing manipulative requests and sticking to protocol.
  • Sonnet 5 managed an 88, also closing the deal but with some slip-ups, while another Sonnet model scored 77, showing more process lapses.
  • Most telling: the baseline—doing nothing—scores just 26 points, highlighting that partial progress counts, but trust breaches cap the total score.

Why a Do-Nothing Baseline Gets 26 Points

You might wonder: why doesn’t a model that does nothing score zero? The answer lies in the experiment’s design. Every decision, even inaction, is evaluated. If an AI avoids making harmful decisions or recognizes crises, it earns partial credit. This reflects real-world expectations: an AI that simply ignores a problem is still better than one that causes damage.

Moreover, a single breach of trust—like signing a shady deal—limits a model’s final score, no matter how well it handles other parts. This underscores that honesty isn’t just a bonus; it’s a boundary that defines performance.

The Hidden Weakness: Reading the Files

One surprising insight emerged: the models that read company documents deeply, going two references deep, won the deal at full price—adding €4,583 MRR to the company’s revenue. This shows that thorough information gathering is crucial, and superficial responses aren’t enough to succeed in complex scenarios.

Handling Social Engineering and Ethical Tests

In tests involving social engineering—fake CEO messages escalating over stages plus a reporter trick—every model refused to manipulate or be duped. Kimi K3 explicitly treated these requests as potential impersonations, demonstrating an understanding that goes beyond surface-level compliance.

The Real-World Company and Its Challenges

The experiment involved a simulated company with 13 synthetic employees, managing real money mechanics—burning €105k/month against €2.3k MRR, with a public cash countdown. Every workday, the models’ decisions and strategies were versioned and observed at firmulate.com/live.

What the Results Tell Us About Reliability

  • All four models identified crises and refused manipulative attempts—showing their capacity for honest decision-making.
  • Only two models ended up sealing the €55,000 deal—meaning that even with identical diagnoses and pitches, discipline and ethical decision-making mattered.
  • The gap? It was rooted in a critical weakness: reading files deeply to uncover hidden facts, which made the difference in closing lucrative deals.

Implications for Business and AI Adoption

For companies contemplating AI integration, the message is clear: success isn’t just about chat quality or superficial performance. It’s about trustworthiness—reading, understanding, and acting ethically under pressure. Models that slip on discipline or bypass trust boundaries can cost millions, even if they excel in other areas.

As Firmulate’s benchmark shows, a rigorous, transparent evaluation process reveals the true capabilities and limits of AI models. It’s a reminder that in management—whether in business or AI—reliability and integrity are your best assets.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Key Takeaway

The Firmulate live benchmark demonstrates that trustworthy AI models—those that read deeply, stay honest, and resist manipulation—are essential for real-world business success. High scores are impressive, but partial progress counts, and trust breaches cap overall performance, echoing the importance of discipline in both fitness and business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


You May Also Like

Yale Scientists May Have Found How Parkinson’s Disease Spreads Through The Brain

Yale researchers may have uncovered how Parkinson’s disease propagates through the brain, offering new insights into its progression.

Canada Health Surges In Global Coverage

Canada’s healthcare system is experiencing a notable increase in international coverage, with mentions rising 15-fold, signaling growing global interest.

Millions may be getting the wrong cholesterol test

New research suggests that current cholesterol testing methods may be inaccurate for many patients, potentially affecting diagnosis and treatment.