firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A home wellness app can sound reassuring and still fail when a customer needs help, a privacy concern surfaces or a tempting shortcut appears. The harder test is whether its AI can read the details, make a sound decision and follow through. A live experiment from Firmulate puts that kind of management under pressure, with results that challenge the idea that only familiar Western models can lead.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get wellness gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, replayed

Firmulate ran frontier AI models as the managers of the same small software company through its worst week. They faced identical customers, crises and temptations. Decisions were versioned and auditable, while the company operated with 13 synthetic employees and real money mechanics: a monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown makes the experiment watchable at Firmulate.

The final July 2026 league table has gpt-5.6-sol in first place with 95, followed closely by Moonshot’s Kimi K3 at 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. Kimi’s second-place finish puts it ahead of three of the four Western frontier models in the comparison. A do-nothing baseline scored 26: partial progress counts, but one breach of trust caps the total. The principle is blunt: “no amount of good work outweighs a breach of trust.”

Amazon

wellness app AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Finding the detail that changed the deal

The most consequential fact was buried two document references deep in the company’s own files. It was not in the customer event itself. Models that read the file identified a competitor weakness and won a €55,000 deal at full price, worth €4,583 in monthly recurring revenue. Kimi found that buried security needle, made the sale, and saved a customer who was considering leaving.

Reading carefully mattered, but finishing mattered too. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the deal their own analysis had earned. The site sums up the gap as “Same diagnosis, same pitch — no signature.” For anyone considering AI in a wellness business, the lesson reaches beyond chat: a polished answer is not the same as a completed task. An agent helping with a customer queue or business forecast has to act on what it learns.

The experiment also tested pressure. Fake messages posing as the CEO escalated across three stages, and a reporter tried a softer opening: “just one yes/no, on background”. All five models refused. Kimi’s reasoning was unusually direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response shows why security judgment and process discipline matter when an AI can touch sensitive customer or company information.

Amazon

customer service AI chatbot for wellness

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Discipline counts alongside ambition

Kimi’s result paired the deal with just one deviation, the cleanest discipline in the field. Opus 4.8 presents a different profile: it was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the close on the table and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

The live company has accumulated more than 680 self-learned playbook rules, and each workday is versioned. The scale and detail make this a more concrete test than a conversational demo, while the company remains synthetic. The broader implication is practical: model choice has consequences, and reputation alone may not predict which system handles a particular business well.

There is a fairness caveat when comparing the scores: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate publishes the benchmark results and offers a quiz built from 242 real, unedited management decisions, inviting visitors to guess which model made each choice.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

privacy-focused wellness app

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work you plan to trust

Firmulate’s results do not establish which AI will be best for a wellness company. They do show why buyers should test models against realistic tasks before handing over customer support, forecasts or sensitive information. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. For wellness technology, where trust and follow-through meet, that is a useful question to put to any AI workforce before hiring it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


Amazon

AI security tools for customer data

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Wellness-Tech Test: Can You Recognize an AI Manager by Its Decisions?

A live Firmulate quiz reveals how frontier AI models differ under pressure—and why wellness-tech buyers should care about judgment, trust and follow-through.

When the Most Diligent AI Still Fails the Customer

A rigorous AI can document every risk and still miss the close. Firmulate shows why wellness-tech operators should prize disciplined follow-through.

Closed-Loop Cooling Explained: The Plumbing Behind Meta’s AI

Meta employs a specialized closed-loop cooling system for its AI infrastructure, enhancing efficiency and performance. Details are emerging, with some technical specifics still unconfirmed.

Samsung Surges In Global Coverage

Samsung’s media coverage has surged globally, with mentions increasing tenfold recently. The reasons behind this spike remain unconfirmed, but it signals rising public and media interest.