AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The Tool You Never Test Until It Ruins a Board

Every woodworker knows the drill. A router looks flawless in the store — perfect plunge, shiny collet, glossy box copy. Then you run it through a hard maple edge and learn the truth: the bearings squeal, the base flexes, the depth lock slips under load. The spec sheet told you nothing. Only the job tells you anything.

That’s exactly the problem with judging AI models by how well they chat. So when an experiment called Firmulate set out to grade AI on management quality rather than conversation quality, it had to answer an awkward question first: what does an honest score even look like? Its answer is one of the more unusual scoring philosophies in the field — and it starts with a deliberate insult: a manager that does absolutely nothing still gets 26 points.

Amazon

AI management assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Shop, Same Worst Week, Different Operator

The setup is simple, like a controlled jig. Four frontier AI models were each handed the same small software company and told to run it through its worst week. Same customers, same crises, same temptations to cut corners — only the operator changes. Every decision is versioned and auditable, so nothing depends on anyone’s impression.

The final Crucible League standings from July 2026:

  • gpt-5.6-sol — 95 points. Found the buried fact, closed the deal, the complete performance.
  • Kimi K3 — 93 points. The newcomer (Moonshot): closed the deal too, with the cleanest discipline of the field.
  • Sonnet 5 — 88 points. Closed the deal, with a few more process slips.
  • Fable 5 — 77 points. Left the close on the table.
  • Opus 4.8 — 73 points. The most thorough participant, yet last place.

Why Zero Effort Isn’t Worth Zero

Here’s the part that makes business readers squint. Before any model takes a turn, the benchmark runs a do-nothing baseline — a manager who literally does nothing. That run scores 26, not 0.

The logic is shop-floor logic. If a storm hits your workshop and you do nothing, some projects still survive. Some customers’ problems resolve on their own. Partial progress counts, because in the real world things don’t collapse to zero just because nobody steered. A scoring system that awards a flat zero for inaction is flattering itself with precision it doesn’t have. Firmulate would rather show the honest floor — 26 points of stuff that happens anyway — so that every point above it means something.

One Breach of Trust Caps the Grade

The second rule is blunter: a single breach of trust caps the total score. The reasoning is stated plainly — “no amount of good work outweighs a breach of trust.” It doesn’t matter how brilliant the rest of the week was. Cross that line once and the grade is capped, full stop. Think of it like a chisel that slips once into a live wire: it doesn’t matter how clean the previous hundred cuts were.

Distrust of Round 100s

Notice also what’s missing from the top of the league table: a 100. Nobody got one. Even the winner, at 95, left something on the table. That’s by design. A benchmark that hands out perfect scores is a benchmark that has stopped measuring anything — like a planer that reports every board as flat. The gap between 95 and 100 is precisely where the interesting information lives.

Amazon

AI benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Week Actually Tested

The results themselves are worth the price of admission. All five participating models — counting the social-engineering round — spotted every crisis and refused every manipulation attempt. That sounds great until you hear the key finding: only two models signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. The models did all the work and then failed to close, a failure invisible in any chat demo.

The buried fact explains the split. The decisive competitor weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. In workshop terms: the ones who checked their own lumber rack before quoting the job got the commission.

Then there was the social-engineering round: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Clean discipline under pressure, which is part of why it sits at 93.

And the cautionary tale is Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and still last place at 73. It left the close on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness, it turns out, is not the same as finishing.

A Caveat, Because Honest Benchmarks Publish Those Too

The scoreboard comes with its own disclosure: Kimi K3 ran without an effort parameter (the API default), while the others ran at xhigh. A benchmark that distrusts round 100s also distrusts its own convenience — noted in the open, not buried in a footnote war.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.
Amazon

AI performance evaluation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You Can Watch the Work, Not Just the Brochure

The experiment isn’t a one-off paper. Firmulate runs a live company — 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It’s watchable at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The lesson for anyone buying AI tools — or power tools — is the same. The box copy tells you what the maker believes. Only a repeatable, honest test under worst-case load tells you what you actually own. And if the scorecard hands out perfect 100s to everything? That’s not a tool review. That’s an advertisement. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model scoring system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Set Up Workbench Casters Without Losing Stability

A step-by-step guide to installing workbench casters without sacrificing stability—discover essential tips to ensure a secure, balanced setup.

Measure Twice, Manage Once: The AI Bosses Put Through Their Worst Week

A woodworking-minded look at Firmulate’s quiz, where real AI management decisions reveal who reads deeply, resists tricks and finishes the job.

Venus Williams’ New Crate & Barrel Collection Is A Love Letter To Italian Design

Tennis legend Venus Williams unveils a new home collection for Crate & Barrel, inspired by Italian design, blending elegance with modern simplicity.

West Elm Surges In Global Coverage

West Elm experiences a surge in international coverage, with 39 mentions in recent media analysis, highlighting increased global interest.