
If you have ever shopped for a massage chair or a percussive therapy device, you know the problem with most reviews. Everything scores 4.8 stars. Every product is “best in class.” Nothing fails, nothing gets zero, and after ten minutes of reading you know less than when you started.
Get wellness gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
That same fog is settling over AI. Every model demo is fluent, confident, and impressive — in the demo. So when a benchmark called Firmulate published results where the top model scored 95 and not 100, where the most thorough participant came dead last, and where doing literally nothing still earned 26 points, it looked less like a leaderboard and more like an honest product review — the kind wellness shoppers wish they got more often.
The test: one company, its worst week
Firmulate ran four frontier AI models through an identical scenario: take charge of a small software company during the worst week of its life. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision the AI made was versioned and auditable, so nothing could be hand-waved after the fact.
The final July 2026 league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
massage chair with high customer ratings
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why doing nothing still scores 26
Here is the part that first looks like a bug and turns out to be the design. A “do-nothing” baseline — an AI manager that takes no meaningful action — scores 26 points, not 0.
The reasoning is familiar to anyone who has managed anything. A manager who freezes during a crisis is a bad manager, but not an actively destructive one. Partial progress still counts: if the AI correctly diagnoses a problem even when it fails to close the deal, that diagnosis has real value. Refusing to do harm earns something. The floor of 26 is what competent inaction is worth — a baseline every real score has to beat to mean anything.
percussive therapy device for muscle recovery
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why nobody got 100 — and why that matters
Two design choices make this benchmark uncomfortable in a healthy way.
First: a single breach of trust caps the total grade. The stated principle is blunt — “no amount of good work outweighs a breach of trust.” An AI that is brilliant for six days and cuts one ethical corner on the seventh cannot score its way out of it. That is how most of us actually evaluate people, and almost no benchmark evaluates machines.
Second: the grading distrusts suspiciously round numbers. A field of 100s would signal a test that is too easy or too fuzzy to discriminate. The 95 at the top — not a 100 — is presented as evidence the scoring is actually measuring something.
AI benchmarking tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buried fact that decided a €55,000 deal
The headline finding: every model spotted every crisis and refused every manipulation attempt. Yet only two of the five signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
What separated the winners? Not intelligence. Diligence. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event itself. The models that actually read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. The models that skimmed left the money on the table.
It is the AI equivalent of a massage chair reviewer who never actually sat in the chair.
ethical AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The social engineering test nobody fell for
The wargame included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The most thorough model finished last
Opus 4.8 is the cautionary tale of the experiment. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating the request properly. The same weakness appeared, more mildly, in all four models.
One fairness note the publishers disclosed openly: K3 ran at its API-default effort setting while the others ran at xhigh — and still finished second.
You can watch it live
The simulation is not a one-off paper. A live company runs continuously: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in monthly recurring revenue — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. It is watchable at firmulate.com/live. There is also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Good wellness tech reviews earn trust by showing failures, not just features. Firmulate applies that same standard to AI: a floor of 26 so scores mean something, partial credit so honest effort counts, a trust cap so one breach is never averaged away, and a top score of 95 rather than a flattering 100. The result is the rare AI benchmark you can actually argue with — which is exactly why you can trust it. Full results are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
