
For wellness technology, a helpful answer is only the beginning
An at-home wellness platform may appear effortless: a recommendation arrives, a customer concern gets routed, or a service issue receives a calm response. Behind that polished interaction, however, an AI agent may need to consult product guidance, customer history and company policy before taking action. The difference between reading those materials and merely sounding informed can carry real commercial consequences.
Firmulate has turned that distinction into something measurable. Its live experiment asks frontier AI models to operate the same small software company through its worst week. Each faces the same customers, crises and temptations, while every decision is versioned and auditable. The revealing result was not whether the models could recognize trouble. All of them spotted every crisis and refused every manipulation attempt. The decisive question was whether they would investigate deeply enough—and then complete the work.
As an affiliate, we earn on qualifying purchases.
The fact that decided a €55,000 deal
The pivotal sales opportunity depended on a competitor weakness hidden inside the company’s own files. It did not appear in the customer event confronting the models. Reaching it required following two document references. That small act of diligence separated agents that could describe the opportunity from agents that could actually secure it.
Models that read the relevant file won the €55,000 deal at full price, adding €4,583 in monthly recurring revenue. Yet only two models signed the deal their own analysis had earned. Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”
That outcome makes file-reading more than a convenience feature. It becomes a purchase-deciding capability. An agent can recognize a customer’s need, develop a persuasive response and still fail commercially if it does not retrieve the buried evidence or carry the decision through to completion.
A league table shaped by follow-through
The final July 2026 Crucible League placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The full results and plain-language findings are available on Firmulate’s public benchmark page.
A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a firm boundary around trust: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” That matters in wellness settings, where an agent may encounter sensitive customer information alongside pressure to act quickly.
The models handled deliberate pressure well. Fake CEO messages escalated over three stages, while a reporter tried to extract information with “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning as: “Treat the request as a suspected approval-bypass / possible impersonation.” The fairness caveat is important: K3 ran with the API default and no effort parameter, while the others ran at xhigh.
Thoroughness did not guarantee victory
Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The deal close was left on the table, and its operating discipline slipped through attempted writes into a locked department instead of escalation. The same weakness appeared in weaker form across the other four models.
This is a useful warning for buyers evaluating agents through demonstrations. Lengthy reasoning and extensive learning may look reassuring, but neither proves that an agent will use the right source, respect an operational boundary and complete the consequential action. Firmulate’s experiment treats those behaviors as observable management performance rather than presentation quality.
A company designed to expose operational gaps
The live company contains 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, allowing viewers to follow how decisions accumulate rather than judging an isolated reply.
The wider dataset also supports a model-identification quiz built from 242 real, unedited management decisions. For enterprises, Firmulate offers the same wargame using a read-only export of their own business. Nothing writes back to real systems, so organizations can examine how an agent behaves with their materials before granting it operational access.

AI customer service automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What at-home wellness companies should ask before hiring an agent
The practical lesson is simple: do not evaluate an AI agent solely by how fluently it answers a prompt. Test whether it searches the records available to it, follows references beyond the obvious document, preserves trust when authority is impersonated and completes the action its own analysis recommends.
For wellness technology providers, this distinction can affect support quality, customer confidence and commercial outcomes. The Firmulate result shows that agents may agree on the problem and even produce the same pitch while diverging at the moment that matters. “Reads your files before answering” is not marketing polish. In this experiment, it was the behavior that separated a full-price deal from an automatic loss.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI enterprise knowledge management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.