
Every woodworker knows the rule: never make the first cut on the good stock. You test the blade on scrap, check the fence, see how it handles the knots — because a mistake on the workpiece can’t be undone. It’s strange, then, that companies are wiring AI agents straight into their CRM, support queues, and forecasts without a single test cut. They judge the tool by its showroom finish — how nicely it chats — instead of how it behaves when the board warps.
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A public experiment called Firmulate has been running exactly the kind of stress test a careful maker would demand: frontier AI models put through a company’s worst week, under load, with real temptations to cut corners. The results are watchable, versioned, and surprisingly humbling.
The Wargame
Firmulate gave four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, like a shop notebook you can replay page by page. The final league from July 2026: gpt-5.6-sol finished first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. For context, doing nothing at all scores 26 — and a single breach of trust caps the total, no matter how good the rest of the work is.
Same Diagnosis, No Signature
Here’s the finding that should stop any manager mid-cut. All four models spotted every crisis. All four refused every manipulation attempt. But only two of them actually closed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in a chat demo. It only shows up under load, on the workpiece.
And the buried fact is better still: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. Diligence paid; surface work didn’t.
Pressure Testing Against Impersonation
The week also included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, every time. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you want proven, not promised.
The Thoroughness Trap
Opus 4.8 is the cautionary tale for anyone who equates effort with results. It was the most thorough participant — 80-plus learned rules, the deepest analyses — and still finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. A bench plane that takes the finest shaving still ruins the piece if you never stop planing.
Watch It Live
Firmulate isn’t a one-off benchmark. There’s a live company running right now — 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com. There’s also a quiz built on 242 real, unedited management decisions where you guess which model made which call — a fast way to feel how differently these tools behave under pressure.
One fairness note: Kimi K3 ran without an effort parameter while the others ran at maximum effort, which makes its second-place finish even more striking.

The lesson transfers straight from the shop: judge a tool by its cuts under load, not by its polish in the catalog. If an AI agent will ever touch your customer data, your pipeline, or your forecast, run it against a crisis first — on scrap, not on the good lumber.
That’s what the Firmulate pilot does for enterprises. You provide a read-only export of your own business — customers, pipeline, rules — and the same wargame runs against it: churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks, and nothing ever writes back to your real systems. Same principle as the test cut: full pressure, zero risk to the workpiece.
Ready to stress-test before you commit the good stock? Start your pilot at firmulate.com/pilot.html or write to contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
