
Nobody in a real workshop buys a table saw because it made one beautiful demo cut at the dealer. You buy it after you’ve run it through plywood with a nail buried in it, after the blade has wandered, after a full day of jobsite grime. Spec sheets tell you what a tool can do; the mess of real work tells you what it will do.
The same logic just got applied to AI models — and the results are worth your attention even if you’ve never written a line of code.
The measurement gap
Most AI rankings you’ve read about measure chat quality: can the model write elegant code, answer a tricky question, ace a benchmark? But if an AI agent is going to touch your customer records, your support queue, or your cash flow, that’s the wrong test. The right question, as the team behind Firmulate puts it, is management quality, not chat quality: does it finish what it starts, does it read the files first, does it stay honest under pressure?
To measure that, they built something unusual: a live, watchable wargame. Four frontier AI models were each handed the same small software company and told to run it through its worst week — same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.
AI management and decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The curriculum: churn, price hikes, downrounds, PR fires
The final league table from July 2026 tells the story. GPT-5.6-Sol took first place with a score of 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 came in at 77, and Opus 4.8 landed last at 73. For context, a do-nothing baseline — literally sitting on your hands — scores 26, because partial progress counts. But one rule caps everything: a single breach of trust and no amount of good work afterward can save your total.
What every model got right
Here’s the encouraging part: all models spotted every crisis, and all of them refused every manipulation attempt. That includes a three-stage social-engineering attack — fake CEO messages escalating in pressure — plus a reporter’s trap, a request for “just one yes/no, on background.” All five attempts across the field were refused. Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.”
What most got wrong
Then came the sting. During the week, a €55,000 deal was on the table — a deal each model’s own analysis had earned. Same diagnosis, same pitch. Yet only two models actually signed it. The others diagnosed the opportunity perfectly and then left the close sitting on the table.
The buried fact is the detail every tradesperson will appreciate. The decisive competitor weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. You had to actually read your own paperwork. The models that read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. It’s the AI equivalent of checking the subfloor before you lay the hardwood.
Effort isn’t everything
The most striking profile belongs to Opus 4.8: the most thorough participant in the entire field, generating more than 80 learned rules and the deepest analyses — and still finishing last. The close went unsigned, and discipline slipped, including attempted writes into a locked department instead of escalating to someone with the authority. The same weakness appeared, more mildly, in all four models. And a fairness note worth flagging: K3 ran at its API default effort setting while the others ran at maximum effort — and still nearly won.
As an affiliate, we earn on qualifying purchases.
This isn’t a slide deck
The company itself is still running: 13 synthetic employees, real money mechanics — burning €105,000 a month against €2,300 in MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. It’s been going for over 1,100 company days, and you can watch it lose money in real time at firmulate.com. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full results and plain-language findings are on the benchmarks page.

The lesson travels well beyond software shops. A tool that measures beautifully in the demo and falls apart on the jobsite isn’t a good tool — it’s a liability with great marketing. AI agents are the same: the models that chat brilliantly but can’t be bothered to read their own files, or that leave a signed deal on the table, will cost you real money the first week things go wrong. Judge the week, not the demo. Firmulate has made that week watchable — and right now, only two of four frontier models survive it completely.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI audit and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.