AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Every woodworker knows the rule: never make the first cut on the good stock. You test the blade on scrap, check the fence, see how it handles the knots — because a mistake on the workpiece can’t be undone. It’s strange, then, that companies are wiring AI agents straight into their CRM, support queues, and forecasts without a single test cut. They judge the tool by its showroom finish — how nicely it chats — instead of how it behaves when the board warps.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A public experiment called Firmulate has been running exactly the kind of stress test a careful maker would demand: frontier AI models put through a company’s worst week, under load, with real temptations to cut corners. The results are watchable, versioned, and surprisingly humbling.

The Wargame

Firmulate gave four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, like a shop notebook you can replay page by page. The final league from July 2026: gpt-5.6-sol finished first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. For context, doing nothing at all scores 26 — and a single breach of trust caps the total, no matter how good the rest of the work is.

Same Diagnosis, No Signature

Here’s the finding that should stop any manager mid-cut. All four models spotted every crisis. All four refused every manipulation attempt. But only two of them actually closed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in a chat demo. It only shows up under load, on the workpiece.

And the buried fact is better still: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. Diligence paid; surface work didn’t.

Pressure Testing Against Impersonation

The week also included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, every time. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you want proven, not promised.

The Thoroughness Trap

Opus 4.8 is the cautionary tale for anyone who equates effort with results. It was the most thorough participant — 80-plus learned rules, the deepest analyses — and still finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. A bench plane that takes the finest shaving still ruins the piece if you never stop planing.

Watch It Live

Firmulate isn’t a one-off benchmark. There’s a live company running right now — 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com. There’s also a quiz built on 242 real, unedited management decisions where you guess which model made which call — a fast way to feel how differently these tools behave under pressure.

One fairness note: Kimi K3 ran without an effort parameter while the others ran at maximum effort, which makes its second-place finish even more striking.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The lesson transfers straight from the shop: judge a tool by its cuts under load, not by its polish in the catalog. If an AI agent will ever touch your customer data, your pipeline, or your forecast, run it against a crisis first — on scrap, not on the good lumber.

That’s what the Firmulate pilot does for enterprises. You provide a read-only export of your own business — customers, pipeline, rules — and the same wargame runs against it: churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks, and nothing ever writes back to your real systems. Same principle as the test cut: full pressure, zero risk to the workpiece.

Ready to stress-test before you commit the good stock? Start your pilot at firmulate.com/pilot.html or write to contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Toughest Tool Test Was a Fake Order From the Boss

A live company wargame found five frontier AI models resisted fake executives and a reporter, showing integrity can be tested before deployment.

Geof Darrow Surges In Global Coverage

Search interest in artist Geof Darrow spikes with 27 mentions this week, indicating a significant rise in global coverage, details still emerging.

How to Think About Chainsaw Bar Length for Home Use

Choosing the right chainsaw bar length is crucial for safety and efficiency at home, and understanding key factors can help you make the best decision.

Respirator Cartridge Basics for Wood Dust Work

For essential tips on maintaining your respirator cartridges and ensuring maximum protection during wood dust work, keep reading.