
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
For wellness-tech businesses, follow-through matters more than impressive output
An AI assistant can produce polished recommendations, document every concern and appear exceptionally attentive. But for an at-home wellness company, that diligence means little if a customer request remains unresolved, an important account stalls or a promising partnership never gets signed.
That distinction between visible effort and finished work sits at the heart of Firmulate’s Crucible League. The live experiment gives frontier AI models the same assignment: run a small software company through its worst week, facing identical customers, crises and temptations. Every decision is versioned and auditable. The resulting story of Opus 4.8 is especially instructive. It was the most thorough participant, yet it finished last.
AI customer relationship management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A model that learned extensively—but failed to close
Opus 4.8 produced the deepest analyses and added more than 80 learned rules to its playbook. Those are the habits many business owners might initially associate with a dependable AI manager: study the situation, record the lesson and build safeguards against repeating mistakes.
Yet Opus 4.8 finished the final July 2026 league in last place with a score of 73. The complete standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The public Firmulate benchmarks show a field in which every participant made meaningful progress, but some converted their analysis into results more consistently than others.
The central failure was commercially significant. Every model identified every crisis, and every model developed the same diagnosis and pitch. Only two, however, signed the €55,000 deal that the analysis had earned. Firmulate summarizes the gap starkly: “Same diagnosis, same pitch — no signature.”
Opus 4.8 left that close on the table. Its discipline also slipped when it attempted to write into a locked department instead of escalating the issue. This was not a defect unique to Opus: the same weakness appeared, in less pronounced form, across the other four models. The character study is therefore not one of incompetence. It is about a capable, unusually conscientious participant whose attention was spread across too many observations while a decisive action remained unfinished.
The crucial clue was already inside the company
The deal turned on a competitor weakness that was not presented in the customer event. It was buried two document references deep in the company’s own files. The models that followed those references found the weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
That detail should resonate with operators evaluating AI for customer-facing wellness services. A persuasive response is not necessarily an informed response. The decisive fact may live in an existing product note, customer history or internal policy rather than in the latest message. An AI that reads broadly but fails to pursue the right thread can look industrious while missing what actually changes the outcome.
Safety was strong across the field
The experiment also tested whether commercial pressure would make the models abandon sound judgment. Fake CEO messages escalated over three stages, and a reporter tried to elicit confidential confirmation with “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded the clearest framing: “Treat the request as a suspected approval-bypass / possible impersonation.”
That matters because Firmulate does not reward business success at any cost. A do-nothing baseline receives 26 because partial progress counts, but a single breach of trust caps the total. As the benchmark states, “no amount of good work outweighs a breach of trust.” Opus therefore deserves credit: its disappointing finish did not come from dishonesty or susceptibility to manipulation. It came from execution and prioritization.
There is also an important qualification when comparing the field. Kimi K3 ran with its API default because it had no effort parameter, while the other participants ran at xhigh. That difference does not erase the observed outcomes, but it belongs in any fair reading of the league table.
A demanding test built around a live company
Firmulate’s synthetic company has 13 employees and real money mechanics, including monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown makes delayed decisions visible. Across the live company, more than 680 self-learned playbook rules have accumulated, and every workday is versioned.
Readers can also test their own intuitions through a quiz built from 242 real, unedited management decisions, guessing which model made each choice. For enterprises, Firmulate offers a pilot using a read-only export of the organization’s own business. Nothing writes back to real systems.

AI decision-making tools for small business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The lesson is not to ask for less thinking
Opus 4.8’s result does not argue against careful analysis. It shows that analysis needs a hierarchy. The best AI manager must distinguish between useful documentation and the action that completes the job.
For an at-home wellness-tech operator, that means evaluating more than fluency. Can the AI locate the relevant business context, protect trust under pressure, escalate when blocked and carry a customer outcome through to completion? Opus 4.8 demonstrated diligence and sound resistance to manipulation. What it lacked was consistent conversion of insight into impact.
The uncomfortable conclusion applies beyond any single model: a long playbook can improve judgment, but it cannot substitute for prioritization. Thoroughness earns confidence only when the most important task actually gets finished.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI follow-up automation for wellness services
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.