firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

When Wellness Software Starts Making Decisions

At-home wellness technology often sits close to personal routines, sensitive information and daily habits. As these products become more capable, the important question is no longer simply whether an AI can produce a polished answer. It is whether the system notices trouble, checks the available evidence, protects trust and completes the job.

Firmulate has turned that question into a public experiment—and a surprisingly revealing guessing game. Its guess-the-model quiz draws on 242 real, unedited management decisions made by frontier AI models. Readers see how a model handled a business situation and try to identify it from the response alone.

The attraction is partly playful, but the subject is serious. Put several capable models into identical circumstances and recognizable management personalities begin to emerge: exhaustive analysis, concise discipline, procedural hesitation and different levels of follow-through.

Amazon

AI management decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Worst Week at the Same Company

Firmulate asked each frontier model to run the same small software company through its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable. That makes the exercise closer to a controlled management wargame than a collection of unrelated chatbot demonstrations.

The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The broad result initially looks reassuring. All models spotted every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

The Detail That Separated Analysis From Action

The decisive weakness of a competitor was not sitting in the customer event where a hurried manager might expect to find it. It was buried two document references deep in the company’s own files. Models that followed the trail found the fact and won the deal at full price, worth +€4,583 MRR.

That episode matters beyond sales. A wellness assistant may sound informed while relying only on the information immediately visible in a conversation. Firmulate’s experiment shows why fluent responses and effective management are different tests. The winning behavior required attention to the company’s own records and the discipline to act on what they revealed.

Every Model Held the Trust Boundary

The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimity is notable for any business considering AI access to customer records, support queues or operational tools. In this experiment, the models did not trade integrity for convenience when the pressure increased. K3’s result does carry a fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh.

The Most Thorough Model Did Not Win

Opus 4.8 offers the quiz’s clearest warning against equating volume with performance. It was the most thorough participant, learned +80 rules and produced the deepest analyses, yet finished last in the model field with 73. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other models, though less strongly.

This is where the decisions start to feel like character studies. One model can examine a problem in remarkable depth and still miss the final operational move. Another can communicate a firm security boundary in a short sentence. The quiz asks readers to identify those patterns without the branding attached.

A Company You Can Watch

The test company has 13 synthetic employees and real money mechanics. It burns €105k/month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than a retrospective showcase.

Firmulate also offers enterprises the same type of wargame against a read-only export of their own business. Nothing writes back to real systems. The proposition is straightforward: evaluate an AI workforce under company-specific pressure before giving it operational responsibility.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

wellness tech AI assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Wellness-Tech Buyers Should Look For

For at-home wellness technology, a model’s tone is only the visible surface. The more consequential questions concern whether it reads the relevant material, resists attempts to bypass approval, preserves trust and finishes what it starts.

Firmulate’s results do not reduce those qualities to a single stereotype. The league table shows that models capable of detecting the same crises can still diverge sharply in execution. Thoroughness can coexist with hesitation; strong analysis can stop before a signature; disciplined refusal can be both concise and effective.

The 242-decision quiz makes those differences tangible. Guessing the model is entertaining, but the deeper exercise is learning to recognize the management behavior hidden beneath fluent language. For anyone evaluating AI-powered wellness products, that may be the more useful skill.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


Amazon

AI data security tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI trust verification

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Tiny JPEGs Look Different In Chrome

Exploring the reasons behind the visual differences of small JPEG images in Chrome and what it means for web developers and users.

Pixel Watch 5

Leaked images and details suggest Pixel Watch 5 will feature a larger display, improved hardware, and new health tracking features, with official launch expected soon.

Pixel 11 Ads Show A Mysterious Screen-equipped Tracker That’s Not Pixel Watch 5 Or Fitbit Air [Gallery]

New Pixel 11 ads show a device with a screen that is not the Pixel Watch 5 or Fitbit Air, sparking speculation about a new wearable or tracking device.