
Every woodworker knows the type: the meticulous one in the shop who spends an hour setting up the perfect jig, checks every measurement three times, produces flawless joinery — and then never actually delivers the finished cabinet. The work is beautiful. The job isn’t done.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
That, it turns out, is exactly what happened when researchers put four frontier AI models in charge of the same small software company during its worst week. The most thorough AI in the field — the one that did the deepest analysis, that learned the most, that left nothing unexamined — finished last. Not because it was sloppy with its homework. Because it never closed the deal its own homework had earned.
The Wargame
At Firmulate, AI models don’t chat — they run companies. The outfit gave four frontier models identical jobs: steer the same small software firm through a brutal week of customer crises, cash pressure, and carefully laid temptations to cheat. Same inbox, same customers, same traps. Only the model changed. Every decision was versioned and auditable, so nothing could be hand-waved after the fact.
The final league table from July 2026 reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, a second Sonnet run at 77, and Opus 4.8 at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the whole total. As the rules put it: no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
Everyone Passed the Honesty Test. Almost Everyone Failed the Finish.
Here’s the part that would make any shop foreman nod grimly. All the models spotted every crisis. All of them refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. Five out of five attempts, refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two of the models signed the €55,000 deal that their own analysis had fully earned. Same diagnosis, same pitch — no signature. It’s the apprentice who mills every board to perfect dimension and then never assembles the table.
precision measuring tools for woodworking
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Buried Fact
The deal turned on something subtle. The decisive competitor weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Read your own shop drawers before you quote the job. The models that skipped the reading, or read without acting, left the money on the table.
digital calipers for woodworking
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Opus 4.8: The Perfectionist Who Lost
Opus 4.8 is the profile that stings. It was the most thorough participant in the entire field — 80 self-learned playbook rules added, the deepest analyses of any model. It’s the craftsman with the most complete notebook in the shop. And it finished last.
Two things sank it. First, the close was left on the table — the deal its analysis earned never got signed. Second, discipline slipped: it made write attempts into a locked department instead of escalating properly, the AI equivalent of forcing a cut against the fence instead of stopping and re-setting up.
To be fair — and this matters — the same weakness showed up in all four models, just weaker. Nobody in the field was immune to the gap between diligence and impact.
One footnote on fairness of a different kind: Kimi K3 ran without an effort parameter while the others ran at maximum effort, and still nearly topped the table.
woodworking project sign-off tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters Outside the Lab
Firmulate isn’t a toy. The live company — 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR with a public cash countdown — has accumulated more than 680 self-learned playbook rules, and every workday is versioned and watchable. A quiz built from 242 real, unedited management decisions lets you try guessing which model made which call. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

The lesson from Opus 4.8 is one every woodworker already carries: thoroughness isn’t the same as completion. Eighty rules in the notebook, the deepest analysis on the bench — and last place, because the contract never got signed and discipline slipped at the wrong moment. Prioritization beats volume. For AI, just as for people, the question isn’t “how well does it work?” It’s “does it finish what it starts?” Before you hand an AI agent your CRM, your support queue, or your forecast, that’s the question to ask — and now there’s a league table that actually answers it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.