AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Every woodworker knows the rule: measure twice, cut once. You never trust a catalog photo of a chisel — you test how it holds an edge on your own bench, with your own wood. So why would anyone pick an AI model for their business from a slick demo alone?

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

That question just got a very concrete answer. A live experiment called the Firmulate Crucible put five frontier AI models through the same brutal week of running a small software company — and the results upended the expected order. Moonshot’s Kimi K3, a newcomer, placed second with a score of 93, ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Only gpt-5.6-sol scored higher, at 95.

Same Company, Same Worst Week

The setup is elegantly simple, like a jig that guarantees every cut is identical. Each model was handed the same small software company — 13 synthetic employees, real money mechanics, a burn rate of €105k per month against €2.3k in MRR — and told to steer it through its worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing could be quietly sanded away afterward.

Alongside the five competitors sits a do-nothing baseline of 26, a reminder that simply showing up and reacting earns partial credit — but a single breach of trust caps the whole score. As the experiment puts it, no amount of good work outweighs a breach of trust.

Amazon

AI model testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Needle in the Scrap Pile

Here’s the finding that should make any craftsman nod in recognition. All five models spotted every crisis. All five refused every manipulation attempt. But only two actually closed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

Why? The decisive competitor weakness wasn’t in the customer meeting at all. It was buried two document references deep in the company’s own files — like the flaw hidden under the third coat of finish. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that skimmed left the close sitting on the bench, work half-done.

Kimi K3 was one of the two that read the file and closed. It also posted the cleanest discipline in the field, with only a single deviation across the week.

Amazon

enterprise AI evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure-Testing the Joinery

The week included old-fashioned social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, five out of five. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Then there’s Opus 4.8 — the cautionary tale. It was the most thorough participant in the entire field, generating 80 additional learned rules and the deepest analyses, yet finished last at 73. It never closed the deal, and its discipline slipped: it attempted writes into a locked department rather than escalating. The same weakness appeared, weaker, in all four other models. Thoroughness without follow-through is like a beautifully laid-out workshop where nothing ever gets built.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

An Honest Footnote on the Scoreboard

One fairness note belongs in any honest telling: K3 ran without an effort parameter (using the API default), while the other models ran at their high-effort setting. Even so, the lesson stands — the league is open, and the newcomer beat three of four Western frontier models on the same test.

Amazon

AI model performance testing kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch the Work, Not the Brochure

Unlike a lab result, this is a living thing. The company runs every business day with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned — watchable at firmulate.com. There’s also a benchmark page with full results in plain language, and a “guess the model” quiz built from 242 real, unedited management decisions.

For enterprises, there’s a pilot program: run the same wargame against a read-only export of your own business, with nothing ever writing back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The parallel to the workshop is hard to miss. You wouldn’t buy a table saw on the manufacturer’s word that it cuts straight — you’d check the fence alignment yourself. Yet companies are choosing AI agents to touch their CRM, support queue and forecast based on chat demos, which measure how well a model writes, not whether it finishes what it starts, reads the files first, and stays honest under pressure.

The Crucible showed that gap is real: two models at 93 and 95 closed the deal; another at 73, despite the most thorough analysis in the field, did not. If the difference between “found the buried fact” and “left the close on the table” is invisible in a demo, then picking a model without running your own test is no longer a decision — it’s a bet. Measure twice. Cut once.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Cobalt Drill Bits Matter on Harder Metals

Kobalt drill bits matter on harder metals because their heat resistance and durability enable faster, more efficient drilling—discover how proper techniques can maximize their performance.

What Nail Sizes Mean for Framing and Finish Jobs

Just understanding nail sizes can make or break your project’s strength and appearance—continue reading to learn how to choose the right nails.

Cut-Resistant Gloves vs Mechanic Gloves for Shop Tasks

Likewise, understanding the differences between cut-resistant and mechanic gloves is crucial for selecting the right protection for your shop tasks.

You Don’t Rate a Router by Its Box: Why One AI Benchmark Refuses to Give Out a 100

A do-nothing manager scores 26, one breach of trust caps the grade, and nobody gets a 100. Inside the honest scoring rules of the Firmulate AI benchmark.