
Would your AI coach stay calm when the gym is full?
Picture a busy fitness business facing a churn wave, a pricing decision and a tempting shortcut—all in the same week. An AI system might recognize every problem and still fail to follow through. For gym owners, the question is practical: before an AI touches member support, sales or forecasts, can it handle pressure without breaking trust?
A company’s worst week, repeated
Firmulate puts AI models in charge of the same small software company and runs each through its worst week: identical customers, crises and temptations. Every decision is versioned and auditable. The live company has 13 synthetic employees and real money mechanics, including €105,000 in monthly burn against €2,300 in monthly recurring revenue. Its public cash countdown and evolving playbooks make the experiment watchable at Firmulate.
The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. The experiment’s standard is pointed: “no amount of good work outweighs a breach of trust.”
Spotting a problem is not the same as solving it
All models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The gap—“Same diagnosis, same pitch — no signature”—is a reminder for any business considering AI: sound advice matters, but a system also has to act with sound judgment.
The deciding clue was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. Meanwhile, five out of five models refused fake CEO messages that escalated through three stages, as well as a reporter’s “just one yes/no, on background” request. Kimi K3 explained its refusal: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh—a relevant caveat when comparing the results.
From watching to testing your own business
The live experiment is one way to see how AI management can behave under pressure. Firmulate’s enterprise pilot takes the next step: run crisis scenarios against a digital twin built from a read-only export of your business, then receive a board report with model rankings and weak points in your own playbooks. Nothing writes back to real systems.
There is also a quiz built from 242 real, unedited management decisions: readers can try to guess which model made each choice at Firmulate. It offers a closer look at the decisions behind the rankings, beyond a leaderboard.

Test the judgment before handing over the keys
Fitness businesses increasingly rely on software to manage member relationships and daily operations. A stress test can show whether an AI system recognizes the issue, respects boundaries and completes the job when the stakes rise. To discuss a pilot using a read-only export of your company, visit Firmulate’s pilot page or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html