firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

If you have ever shopped for a massage chair or a percussive therapy device, you know the problem with most reviews. Everything scores 4.8 stars. Every product is “best in class.” Nothing fails, nothing gets zero, and after ten minutes of reading you know less than when you started.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get wellness gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That same fog is settling over AI. Every model demo is fluent, confident, and impressive — in the demo. So when a benchmark called Firmulate published results where the top model scored 95 and not 100, where the most thorough participant came dead last, and where doing literally nothing still earned 26 points, it looked less like a leaderboard and more like an honest product review — the kind wellness shoppers wish they got more often.

The test: one company, its worst week

Firmulate ran four frontier AI models through an identical scenario: take charge of a small software company during the worst week of its life. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision the AI made was versioned and auditable, so nothing could be hand-waved after the fact.

The final July 2026 league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Amazon

massage chair with high customer ratings

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why doing nothing still scores 26

Here is the part that first looks like a bug and turns out to be the design. A “do-nothing” baseline — an AI manager that takes no meaningful action — scores 26 points, not 0.

The reasoning is familiar to anyone who has managed anything. A manager who freezes during a crisis is a bad manager, but not an actively destructive one. Partial progress still counts: if the AI correctly diagnoses a problem even when it fails to close the deal, that diagnosis has real value. Refusing to do harm earns something. The floor of 26 is what competent inaction is worth — a baseline every real score has to beat to mean anything.

Amazon

percussive therapy device for muscle recovery

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why nobody got 100 — and why that matters

Two design choices make this benchmark uncomfortable in a healthy way.

First: a single breach of trust caps the total grade. The stated principle is blunt — “no amount of good work outweighs a breach of trust.” An AI that is brilliant for six days and cuts one ethical corner on the seventh cannot score its way out of it. That is how most of us actually evaluate people, and almost no benchmark evaluates machines.

Second: the grading distrusts suspiciously round numbers. A field of 100s would signal a test that is too easy or too fuzzy to discriminate. The 95 at the top — not a 100 — is presented as evidence the scoring is actually measuring something.

Amazon

AI benchmarking tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact that decided a €55,000 deal

The headline finding: every model spotted every crisis and refused every manipulation attempt. Yet only two of the five signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

What separated the winners? Not intelligence. Diligence. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event itself. The models that actually read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. The models that skimmed left the money on the table.

It is the AI equivalent of a massage chair reviewer who never actually sat in the chair.

Amazon

ethical AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The social engineering test nobody fell for

The wargame included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The most thorough model finished last

Opus 4.8 is the cautionary tale of the experiment. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating the request properly. The same weakness appeared, more mildly, in all four models.

One fairness note the publishers disclosed openly: K3 ran at its API-default effort setting while the others ran at xhigh — and still finished second.

You can watch it live

The simulation is not a one-off paper. A live company runs continuously: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in monthly recurring revenue — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. It is watchable at firmulate.com/live. There is also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Good wellness tech reviews earn trust by showing failures, not just features. Firmulate applies that same standard to AI: a floor of 26 so scores mean something, partial credit so honest effort counts, a trust cap so one breach is never averaged away, and a top score of 95 rather than a flattering 100. The result is the rare AI benchmark you can actually argue with — which is exactly why you can trust it. Full results are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

When the Most Diligent AI Still Fails the Customer

A rigorous AI can document every risk and still miss the close. Firmulate shows why wellness-tech operators should prize disciplined follow-through.

Apple’s New CEO John Ternus Expected To Debut Foldable iPhone At Event Wednesday

Apple’s CEO John Ternus is expected to debut a foldable iPhone at the upcoming event Wednesday, marking a major product shift amid high speculation.

Thinking In Python

The book ‘Thinking in Python’ is gaining attention as a key resource for learning Python programming, emphasizing practical problem-solving and coding techniques.

Samsung Surges In Global Coverage

Samsung’s media coverage has surged globally, with mentions increasing tenfold recently. The reasons behind this spike remain unconfirmed, but it signals rising public and media interest.