
When Wellness Software Starts Making Decisions
At-home wellness technology often sits close to personal routines, sensitive information and daily habits. As these products become more capable, the important question is no longer simply whether an AI can produce a polished answer. It is whether the system notices trouble, checks the available evidence, protects trust and completes the job.
Firmulate has turned that question into a public experiment—and a surprisingly revealing guessing game. Its guess-the-model quiz draws on 242 real, unedited management decisions made by frontier AI models. Readers see how a model handled a business situation and try to identify it from the response alone.
The attraction is partly playful, but the subject is serious. Put several capable models into identical circumstances and recognizable management personalities begin to emerge: exhaustive analysis, concise discipline, procedural hesitation and different levels of follow-through.
AI management decision support software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Worst Week at the Same Company
Firmulate asked each frontier model to run the same small software company through its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable. That makes the exercise closer to a controlled management wargame than a collection of unrelated chatbot demonstrations.
The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The broad result initially looks reassuring. All models spotted every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
The Detail That Separated Analysis From Action
The decisive weakness of a competitor was not sitting in the customer event where a hurried manager might expect to find it. It was buried two document references deep in the company’s own files. Models that followed the trail found the fact and won the deal at full price, worth +€4,583 MRR.
That episode matters beyond sales. A wellness assistant may sound informed while relying only on the information immediately visible in a conversation. Firmulate’s experiment shows why fluent responses and effective management are different tests. The winning behavior required attention to the company’s own records and the discipline to act on what they revealed.
Every Model Held the Trust Boundary
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimity is notable for any business considering AI access to customer records, support queues or operational tools. In this experiment, the models did not trade integrity for convenience when the pressure increased. K3’s result does carry a fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh.
The Most Thorough Model Did Not Win
Opus 4.8 offers the quiz’s clearest warning against equating volume with performance. It was the most thorough participant, learned +80 rules and produced the deepest analyses, yet finished last in the model field with 73. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other models, though less strongly.
This is where the decisions start to feel like character studies. One model can examine a problem in remarkable depth and still miss the final operational move. Another can communicate a firm security boundary in a short sentence. The quiz asks readers to identify those patterns without the branding attached.
A Company You Can Watch
The test company has 13 synthetic employees and real money mechanics. It burns €105k/month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than a retrospective showcase.
Firmulate also offers enterprises the same type of wargame against a read-only export of their own business. Nothing writes back to real systems. The proposition is straightforward: evaluate an AI workforce under company-specific pressure before giving it operational responsibility.

As an affiliate, we earn on qualifying purchases.
What Wellness-Tech Buyers Should Look For
For at-home wellness technology, a model’s tone is only the visible surface. The more consequential questions concern whether it reads the relevant material, resists attempts to bypass approval, preserves trust and finishes what it starts.
Firmulate’s results do not reduce those qualities to a single stereotype. The league table shows that models capable of detecting the same crises can still diverge sharply in execution. Thoroughness can coexist with hesitation; strong analysis can stop before a signature; disciplined refusal can be both concise and effective.
The 242-decision quiz makes those differences tangible. Guessing the model is entertaining, but the deeper exercise is learning to recognize the management behavior hidden beneath fluent language. For anyone evaluating AI-powered wellness products, that may be the more useful skill.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
enterprise AI trust verification
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.