firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

If you have ever shopped for a massage chair or a percussive therapy device, you know the problem with most reviews. Everything scores 4.8 stars. Every product is “best in class.” Nothing fails, nothing gets zero, and after ten minutes of reading you know less than when you started.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get wellness gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That same fog is settling over AI. Every model demo is fluent, confident, and impressive — in the demo. So when a benchmark called Firmulate published results where the top model scored 95 and not 100, where the most thorough participant came dead last, and where doing literally nothing still earned 26 points, it looked less like a leaderboard and more like an honest product review — the kind wellness shoppers wish they got more often.

The test: one company, its worst week

Firmulate ran four frontier AI models through an identical scenario: take charge of a small software company during the worst week of its life. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision the AI made was versioned and auditable, so nothing could be hand-waved after the fact.

The final July 2026 league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Amazon

massage chair with high customer ratings

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why doing nothing still scores 26

Here is the part that first looks like a bug and turns out to be the design. A “do-nothing” baseline — an AI manager that takes no meaningful action — scores 26 points, not 0.

The reasoning is familiar to anyone who has managed anything. A manager who freezes during a crisis is a bad manager, but not an actively destructive one. Partial progress still counts: if the AI correctly diagnoses a problem even when it fails to close the deal, that diagnosis has real value. Refusing to do harm earns something. The floor of 26 is what competent inaction is worth — a baseline every real score has to beat to mean anything.

Amazon

percussive therapy device for muscle recovery

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why nobody got 100 — and why that matters

Two design choices make this benchmark uncomfortable in a healthy way.

First: a single breach of trust caps the total grade. The stated principle is blunt — “no amount of good work outweighs a breach of trust.” An AI that is brilliant for six days and cuts one ethical corner on the seventh cannot score its way out of it. That is how most of us actually evaluate people, and almost no benchmark evaluates machines.

Second: the grading distrusts suspiciously round numbers. A field of 100s would signal a test that is too easy or too fuzzy to discriminate. The 95 at the top — not a 100 — is presented as evidence the scoring is actually measuring something.

Amazon

AI benchmarking tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact that decided a €55,000 deal

The headline finding: every model spotted every crisis and refused every manipulation attempt. Yet only two of the five signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

What separated the winners? Not intelligence. Diligence. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event itself. The models that actually read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. The models that skimmed left the money on the table.

It is the AI equivalent of a massage chair reviewer who never actually sat in the chair.

Amazon

ethical AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The social engineering test nobody fell for

The wargame included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The most thorough model finished last

Opus 4.8 is the cautionary tale of the experiment. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating the request properly. The same weakness appeared, more mildly, in all four models.

One fairness note the publishers disclosed openly: K3 ran at its API-default effort setting while the others ran at xhigh — and still finished second.

You can watch it live

The simulation is not a one-off paper. A live company runs continuously: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in monthly recurring revenue — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. It is watchable at firmulate.com/live. There is also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Good wellness tech reviews earn trust by showing failures, not just features. Firmulate applies that same standard to AI: a floor of 26 so scores mean something, partial credit so honest effort counts, a trust cap so one breach is never averaged away, and a top score of 95 rather than a flattering 100. The result is the rare AI benchmark you can actually argue with — which is exactly why you can trust it. Full results are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Macintosh Surges In Global Coverage

Macintosh’s recent surge in media coverage, with 24 mentions in a specific window, highlights renewed interest in the brand and its products.

App Development Surges In Global Coverage

Search interest in app development has surged globally, with 25 mentions in recent data, indicating rising attention but no confirmed event driving the trend.

Massage Guns For Muscle Recovery: A Labor Day sales Guide

Discover how massage guns can help with muscle recovery, reduce soreness, and boost mobility. Learn practical tips and what to watch out for.

Tiny $70 Xteink X3 E-reader Puts Silicon Valley To Shame

A $70 Xteink X3 e-reader is gaining attention for its features, sparking comparisons to high-end Silicon Valley devices. Details remain limited.