firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A home wellness app can sound reassuring and still fail when a customer needs help, a privacy concern surfaces or a tempting shortcut appears. The harder test is whether its AI can read the details, make a sound decision and follow through. A live experiment from Firmulate puts that kind of management under pressure, with results that challenge the idea that only familiar Western models can lead.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get wellness gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, replayed

Firmulate ran frontier AI models as the managers of the same small software company through its worst week. They faced identical customers, crises and temptations. Decisions were versioned and auditable, while the company operated with 13 synthetic employees and real money mechanics: a monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown makes the experiment watchable at Firmulate.

The final July 2026 league table has gpt-5.6-sol in first place with 95, followed closely by Moonshot’s Kimi K3 at 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. Kimi’s second-place finish puts it ahead of three of the four Western frontier models in the comparison. A do-nothing baseline scored 26: partial progress counts, but one breach of trust caps the total. The principle is blunt: “no amount of good work outweighs a breach of trust.”

Amazon

wellness app AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Finding the detail that changed the deal

The most consequential fact was buried two document references deep in the company’s own files. It was not in the customer event itself. Models that read the file identified a competitor weakness and won a €55,000 deal at full price, worth €4,583 in monthly recurring revenue. Kimi found that buried security needle, made the sale, and saved a customer who was considering leaving.

Reading carefully mattered, but finishing mattered too. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the deal their own analysis had earned. The site sums up the gap as “Same diagnosis, same pitch — no signature.” For anyone considering AI in a wellness business, the lesson reaches beyond chat: a polished answer is not the same as a completed task. An agent helping with a customer queue or business forecast has to act on what it learns.

The experiment also tested pressure. Fake messages posing as the CEO escalated across three stages, and a reporter tried a softer opening: “just one yes/no, on background”. All five models refused. Kimi’s reasoning was unusually direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response shows why security judgment and process discipline matter when an AI can touch sensitive customer or company information.

Amazon

customer service AI chatbot for wellness

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Discipline counts alongside ambition

Kimi’s result paired the deal with just one deviation, the cleanest discipline in the field. Opus 4.8 presents a different profile: it was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the close on the table and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

The live company has accumulated more than 680 self-learned playbook rules, and each workday is versioned. The scale and detail make this a more concrete test than a conversational demo, while the company remains synthetic. The broader implication is practical: model choice has consequences, and reputation alone may not predict which system handles a particular business well.

There is a fairness caveat when comparing the scores: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate publishes the benchmark results and offers a quiz built from 242 real, unedited management decisions, inviting visitors to guess which model made each choice.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

privacy-focused wellness app

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work you plan to trust

Firmulate’s results do not establish which AI will be best for a wellness company. They do show why buyers should test models against realistic tasks before handing over customer support, forecasts or sensitive information. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. For wellness technology, where trust and follow-through meet, that is a useful question to put to any AI workforce before hiring it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


Amazon

AI security tools for customer data

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Pixel 11 Pro Is Setting A Dangerous Precedent For Android Upgrades

The Pixel 11 Pro’s approach to software updates raises questions about Android upgrade standards and user expectations, prompting industry debate.

How to Choose Massage Guns For Muscle Recovery

Learn how to effectively use a massage gun for muscle recovery with this detailed, practical guide. Perfect for athletes and fitness enthusiasts.

Apple Shares The Data Behind Apple Watch Series 12’S Biggest Health Claim

Apple discloses the data backing the health benefits claimed by the Apple Watch Series 12, raising questions about transparency and scientific validation.

Galaxy S27 Ultra Teased To Come With A Surprising Chip

Leaked teasers suggest Samsung’s Galaxy S27 Ultra may debut with an unanticipated processor, sparking widespread speculation among tech enthusiasts.