
Every woodworker knows the rule: measure twice, cut once. You never trust a catalog photo of a chisel — you test how it holds an edge on your own bench, with your own wood. So why would anyone pick an AI model for their business from a slick demo alone?
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
That question just got a very concrete answer. A live experiment called the Firmulate Crucible put five frontier AI models through the same brutal week of running a small software company — and the results upended the expected order. Moonshot’s Kimi K3, a newcomer, placed second with a score of 93, ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Only gpt-5.6-sol scored higher, at 95.
Same Company, Same Worst Week
The setup is elegantly simple, like a jig that guarantees every cut is identical. Each model was handed the same small software company — 13 synthetic employees, real money mechanics, a burn rate of €105k per month against €2.3k in MRR — and told to steer it through its worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing could be quietly sanded away afterward.
Alongside the five competitors sits a do-nothing baseline of 26, a reminder that simply showing up and reacting earns partial credit — but a single breach of trust caps the whole score. As the experiment puts it, no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
The Needle in the Scrap Pile
Here’s the finding that should make any craftsman nod in recognition. All five models spotted every crisis. All five refused every manipulation attempt. But only two actually closed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
Why? The decisive competitor weakness wasn’t in the customer meeting at all. It was buried two document references deep in the company’s own files — like the flaw hidden under the third coat of finish. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that skimmed left the close sitting on the bench, work half-done.
Kimi K3 was one of the two that read the file and closed. It also posted the cleanest discipline in the field, with only a single deviation across the week.
As an affiliate, we earn on qualifying purchases.
Pressure-Testing the Joinery
The week included old-fashioned social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, five out of five. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then there’s Opus 4.8 — the cautionary tale. It was the most thorough participant in the entire field, generating 80 additional learned rules and the deepest analyses, yet finished last at 73. It never closed the deal, and its discipline slipped: it attempted writes into a locked department rather than escalating. The same weakness appeared, weaker, in all four other models. Thoroughness without follow-through is like a beautifully laid-out workshop where nothing ever gets built.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
An Honest Footnote on the Scoreboard
One fairness note belongs in any honest telling: K3 ran without an effort parameter (using the API default), while the other models ran at their high-effort setting. Even so, the lesson stands — the league is open, and the newcomer beat three of four Western frontier models on the same test.
AI model performance testing kit
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Watch the Work, Not the Brochure
Unlike a lab result, this is a living thing. The company runs every business day with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned — watchable at firmulate.com. There’s also a benchmark page with full results in plain language, and a “guess the model” quiz built from 242 real, unedited management decisions.
For enterprises, there’s a pilot program: run the same wargame against a read-only export of your own business, with nothing ever writing back to real systems.

The parallel to the workshop is hard to miss. You wouldn’t buy a table saw on the manufacturer’s word that it cuts straight — you’d check the fence alignment yourself. Yet companies are choosing AI agents to touch their CRM, support queue and forecast based on chat demos, which measure how well a model writes, not whether it finishes what it starts, reads the files first, and stays honest under pressure.
The Crucible showed that gap is real: two models at 93 and 95 closed the deal; another at 73, despite the most thorough analysis in the field, did not. If the difference between “found the buried fact” and “left the close on the table” is invisible in a demo, then picking a model without running your own test is no longer a decision — it’s a bet. Measure twice. Cut once.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
