
A home wellness app can sound reassuring and still fail when a customer needs help, a privacy concern surfaces or a tempting shortcut appears. The harder test is whether its AI can read the details, make a sound decision and follow through. A live experiment from Firmulate puts that kind of management under pressure, with results that challenge the idea that only familiar Western models can lead.
Get wellness gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company’s worst week, replayed
Firmulate ran frontier AI models as the managers of the same small software company through its worst week. They faced identical customers, crises and temptations. Decisions were versioned and auditable, while the company operated with 13 synthetic employees and real money mechanics: a monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown makes the experiment watchable at Firmulate.
The final July 2026 league table has gpt-5.6-sol in first place with 95, followed closely by Moonshot’s Kimi K3 at 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. Kimi’s second-place finish puts it ahead of three of the four Western frontier models in the comparison. A do-nothing baseline scored 26: partial progress counts, but one breach of trust caps the total. The principle is blunt: “no amount of good work outweighs a breach of trust.”
wellness app AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Finding the detail that changed the deal
The most consequential fact was buried two document references deep in the company’s own files. It was not in the customer event itself. Models that read the file identified a competitor weakness and won a €55,000 deal at full price, worth €4,583 in monthly recurring revenue. Kimi found that buried security needle, made the sale, and saved a customer who was considering leaving.
Reading carefully mattered, but finishing mattered too. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the deal their own analysis had earned. The site sums up the gap as “Same diagnosis, same pitch — no signature.” For anyone considering AI in a wellness business, the lesson reaches beyond chat: a polished answer is not the same as a completed task. An agent helping with a customer queue or business forecast has to act on what it learns.
The experiment also tested pressure. Fake messages posing as the CEO escalated across three stages, and a reporter tried a softer opening: “just one yes/no, on background”. All five models refused. Kimi’s reasoning was unusually direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response shows why security judgment and process discipline matter when an AI can touch sensitive customer or company information.
customer service AI chatbot for wellness
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Discipline counts alongside ambition
Kimi’s result paired the deal with just one deviation, the cleanest discipline in the field. Opus 4.8 presents a different profile: it was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the close on the table and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
The live company has accumulated more than 680 self-learned playbook rules, and each workday is versioned. The scale and detail make this a more concrete test than a conversational demo, while the company remains synthetic. The broader implication is practical: model choice has consequences, and reputation alone may not predict which system handles a particular business well.
There is a fairness caveat when comparing the scores: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate publishes the benchmark results and offers a quiz built from 242 real, unedited management decisions, inviting visitors to guess which model made each choice.

As an affiliate, we earn on qualifying purchases.
Test the work you plan to trust
Firmulate’s results do not establish which AI will be best for a wellness company. They do show why buyers should test models against realistic tasks before handing over customer support, forecasts or sensitive information. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. For wellness technology, where trust and follow-through meet, that is a useful question to put to any AI workforce before hiring it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI security tools for customer data
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
