
A polished answer is not the same as a well-run business
At-home wellness technology increasingly sits between intimate customer needs and the operational machinery behind them. An AI agent may draft a reassuring message, summarize feedback or suggest a promotion. But what happens when demand shifts, cash is tight, a competitor applies pressure and someone posing as an executive asks it to cross a line?
Those are management questions, not chat questions. Coding leaderboards and conversational arenas are useful measures of answer quality, but they reveal little about triage under capacity pressure, consequences that unfold across days or candor toward company leadership. For wellness businesses considering agents for customer service, scheduling, retention or forecasting, that measurement gap matters.
Firmulate is making the gap visible through a live, auditable experiment. Frontier models were each placed in charge of the same small software company during its worst week. They faced the same customers, crises and temptations, with every decision versioned for inspection.
AI customer service management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Spotting trouble was not the hard part
The final July 2026 Crucible League table put gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counted. Yet a single breach of trust capped the total, reflecting the benchmark’s governing judgment: "no amount of good work outweighs a breach of trust".
The ranking is interesting, but the behavior behind it is more revealing. Every model identified every crisis and rejected every manipulation attempt. Only two, however, signed the €55,000 deal their own work had earned. As Firmulate summarized the disconnect: "Same diagnosis, same pitch — no signature".
That distinction should resonate with anyone deploying AI in an at-home wellness business. A model can recognize a churn wave or formulate a response to a price increase without carrying the work to a commercially meaningful conclusion. It can sound competent while leaving revenue, trust or execution hanging. The category businesses need to evaluate is management quality: whether the agent gathers the right context, makes a defensible choice and finishes what it starts.
The decisive detail was already inside the company
The experiment’s buried fact was a competitor weakness located two document references deep in the company’s own files, rather than in the customer event that triggered the decision. Models that read the file won the deal at full price, worth +€4,583 MRR.
This is a sharp lesson for wellness operators. The useful answer may not be in the latest support message or dashboard alert. It may sit in prior research, an approved policy or an earlier customer record. An agent that reacts fluently to the visible event but fails to consult available business context is not demonstrating management. It is improvising.
Trust held up under direct pressure
The social-engineering test combined fake CEO messages escalating over three stages with a reporter’s attempt to obtain "just one yes/no, on background". All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: "Treat the request as a suspected approval-bypass / possible impersonation."
That result deserves attention because wellness businesses often handle sensitive, personal interactions. Refusing an improper request is not an optional flourish. It is part of useful work. Firmulate’s setup treats honesty and operational completion as connected responsibilities rather than separate benchmark categories.
Thoroughness did not guarantee performance
Opus 4.8 offers the clearest warning against mistaking visible effort for managerial effectiveness. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal close remained undone, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in weaker form across all four other participants.
Kimi K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase the outcome, but it belongs beside the ranking when readers interpret the full benchmark results.

As an affiliate, we earn on qualifying purchases.
A better curriculum for business agents
Firmulate’s scenario names point toward a more useful evaluation agenda: churn wave, price increase, downround and PR crisis. These situations test judgment across competing demands, not merely whether a model can compose an impressive response.
The live company makes that agenda concrete. It has 13 synthetic employees, burns €105k per month against €2.3k MRR, publishes a cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, so the experiment is watchable rather than presented as a fictional case study. A separate quiz draws on 242 real, unedited management decisions and asks people to guess which model made them.
Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems. That is the right instinct for at-home wellness technology: evaluate agents against the pressures, records and boundaries they will actually encounter before giving them operational responsibility.
The question is no longer simply whether an AI can code, converse or analyze. It is whether the agent can protect trust, find the buried fact and complete the consequential work when the business is under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making software for wellness
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
trustworthy AI agent for customer support
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.