AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Nobody in a real workshop buys a table saw because it made one beautiful demo cut at the dealer. You buy it after you’ve run it through plywood with a nail buried in it, after the blade has wandered, after a full day of jobsite grime. Spec sheets tell you what a tool can do; the mess of real work tells you what it will do.

The same logic just got applied to AI models — and the results are worth your attention even if you’ve never written a line of code.

The measurement gap

Most AI rankings you’ve read about measure chat quality: can the model write elegant code, answer a tricky question, ace a benchmark? But if an AI agent is going to touch your customer records, your support queue, or your cash flow, that’s the wrong test. The right question, as the team behind Firmulate puts it, is management quality, not chat quality: does it finish what it starts, does it read the files first, does it stay honest under pressure?

To measure that, they built something unusual: a live, watchable wargame. Four frontier AI models were each handed the same small software company and told to run it through its worst week — same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.

Amazon

AI management and decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The curriculum: churn, price hikes, downrounds, PR fires

The final league table from July 2026 tells the story. GPT-5.6-Sol took first place with a score of 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 came in at 77, and Opus 4.8 landed last at 73. For context, a do-nothing baseline — literally sitting on your hands — scores 26, because partial progress counts. But one rule caps everything: a single breach of trust and no amount of good work afterward can save your total.

What every model got right

Here’s the encouraging part: all models spotted every crisis, and all of them refused every manipulation attempt. That includes a three-stage social-engineering attack — fake CEO messages escalating in pressure — plus a reporter’s trap, a request for “just one yes/no, on background.” All five attempts across the field were refused. Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.”

What most got wrong

Then came the sting. During the week, a €55,000 deal was on the table — a deal each model’s own analysis had earned. Same diagnosis, same pitch. Yet only two models actually signed it. The others diagnosed the opportunity perfectly and then left the close sitting on the table.

The buried fact is the detail every tradesperson will appreciate. The decisive competitor weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. You had to actually read your own paperwork. The models that read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. It’s the AI equivalent of checking the subfloor before you lay the hardwood.

Effort isn’t everything

The most striking profile belongs to Opus 4.8: the most thorough participant in the entire field, generating more than 80 learned rules and the deepest analyses — and still finishing last. The close went unsigned, and discipline slipped, including attempted writes into a locked department instead of escalating to someone with the authority. The same weakness appeared, more mildly, in all four models. And a fairness note worth flagging: K3 ran at its API default effort setting while the others ran at maximum effort — and still nearly won.

Amazon

AI model testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

This isn’t a slide deck

The company itself is still running: 13 synthetic employees, real money mechanics — burning €105,000 a month against €2,300 in MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. It’s been going for over 1,100 company days, and you can watch it lose money in real time at firmulate.com. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full results and plain-language findings are on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The lesson travels well beyond software shops. A tool that measures beautifully in the demo and falls apart on the jobsite isn’t a good tool — it’s a liability with great marketing. AI agents are the same: the models that chat brilliantly but can’t be bothered to read their own files, or that leave a signed deal on the table, will cost you real money the first week things go wrong. Judge the week, not the demo. Firmulate has made that week watchable — and right now, only two of four frontier models survive it completely.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI audit and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI performance monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Measure Twice, Manage Once: The AI Bosses Put Through Their Worst Week

A woodworking-minded look at Firmulate’s quiz, where real AI management decisions reveal who reads deeply, resists tricks and finishes the job.

What Makes a Good Stop Block System for Repetitive Cuts

Inevitably, choosing the right stop block system ensures precision and safety, but discovering the key factors can significantly improve your setup.

Why Welding Clamps Save More Time Than You Expect

AIThis post was created with the assistance of artificial intelligence (AI).Welding clamps…

Why Oscillating Tool Blades Wear So Differently

AIThis post was created with the assistance of artificial intelligence (AI).Your oscillating…