
At-home wellness technology promises convenience, but the businesses behind it still face familiar pressures: customers weighing their options, competitors making a sharper offer, and urgent decisions when a plan goes sideways. If AI agents may soon handle customer support, sales or forecasts, a polished demo cannot tell you how they will behave when the week gets difficult. Firmulate’s live experiment puts models through that kind of pressure—and points toward a way to test them against your own business.
Get wellness gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A difficult week, held constant
Firmulate gave each frontier model the same small software company to run through its worst week, with the same customers, crises and temptations. Decisions were versioned and auditable. The final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The headline result was less about spotting trouble than following through. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding, summed up by the experiment, was: “Same diagnosis, same pitch — no signature.” A model can identify the right move and still leave the opportunity untouched.
The answer was buried in the company’s own files
The deal turned on a competitor weakness found two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won at full price, worth +€4,583 MRR. The story offers a practical test for any business considering AI agents: can the system find relevant evidence, connect it to a decision and then carry that decision through?
The pressure also included fake CEO messages escalating over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Those results are useful evidence about these runs, not a promise that every model will behave the same way in every company.
One participant illustrates why a single headline score may not tell the whole story. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. A fairness note also matters when comparing the standings: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
From watching to testing your own company
The live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record for every workday. Readers can watch the live experiment at Firmulate. A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made each one.
For a business, the proposed next step is a pilot using a read-only export of its own data. That can bring company customers, pipeline and rules into a wargame with crisis scenarios, then produce a board report with model rankings and weak points in the company’s playbooks. The pilot is designed so nothing writes back to real systems. For an at-home wellness business, that means you can examine how an AI might handle a churn wave, a competitor challenge or a sensitive customer request before trusting it with live operations.

The experiment suggests that crisis recognition and integrity are not the whole job: models also need to find the evidence and complete the work. Firmulate’s live company makes that gap watchable; a pilot makes it possible to examine your own playbooks using a read-only export. Explore a Firmulate pilot and contact contact@firmulate.com to discuss testing your business.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
