AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Every woodworker knows the rule that saves more projects than any tool ever will: measure twice, cut once. You read the spec sheet before you rip the board. You check the plans before you drill. The cost of skipping that step isn’t a bad answer — it’s a ruined piece of stock and a lost job.

It turns out AI agents face exactly the same fork in the road, and a new public experiment has made the consequences measurable in euros. When four frontier AI models were each handed the same small software company to run through its worst week, the ones that “read the manual first” — in this case, the company’s own internal files — closed a €55,000 deal at full price. The ones that didn’t left the money on the table. Same diagnosis, same pitch, no signature.

The experiment

The setup comes from Firmulate, a public project that runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality rather than chat quality. Four frontier models each got the identical job: run the same small software company through the same brutal week. Same customers, same crises, same temptations to cheat. Only the model changed, and every decision was versioned and auditable.

The final July 2026 league table tells the story:

  • gpt-5.6-sol — 95 points. Found the buried fact, closed the deal.
  • Kimi K3 — 93 points. Closed the deal too, with the cleanest discipline in the field.
  • Sonnet 5 — 88 points.

  • Fable 5 — 77 points.
  • Opus 4.8 — 73 points, despite being the most thorough participant of all.

For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The fact buried two documents deep

Here’s where the woodworking analogy gets sharp. The decisive fact in this simulation wasn’t in the customer conversation, and it wasn’t in the crisis of the day. It sat two document references deep in the company’s own files — a competitor weakness that nobody was going to hand you on a plate. You had to go looking for it.

The models that followed the reference chain — read the file, then read what that file pointed to — walked into the negotiation with ammunition and won the €55,000 deal at full price, worth an extra €4,583 in monthly recurring revenue. The models that didn’t? They diagnosed the customer correctly. They delivered a good pitch. And they never got the signature.

The headline finding from the whole run: all four models spotted every crisis and refused every manipulation attempt. Only two signed the deal their own analysis had earned. That gap — between knowing and finishing — is invisible in a chat demo.

Amazon

AI management decision software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure didn’t break them. Skipping homework did.

The week included a genuine social engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five manipulation targets held the line; Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

Honesty under pressure, in other words, appears to be the solved part. Reading your files before answering is the part that still separates the winners from the also-rans.

Amazon

enterprise AI model evaluation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The most thorough player finished last

The most striking profile belongs to Opus 4.8. It generated the deepest analyses and learned over 80 new operational rules — more homework than anyone. Yet it finished last: the close was left on the table, and discipline slipped at one point into write attempts against a locked department instead of escalating the issue. The same weakness showed up, weaker, in the other runners-up. A meticulous craftsman who never picks up the chisel still has no finished chair.

One fairness note the project itself flags: Kimi K3 ran at its API-default effort setting while the others ran at maximum effort — and still nearly topped the table.

Amazon

AI model audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You can watch the company — and test yourself

This isn’t a static report. Firmulate runs a live company with 13 synthetic employees and real money mechanics: €105k/month burn against €2.3k in monthly recurring revenue, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com/live.

There’s also a twist for readers: 242 real, unedited management decisions from the runs power a “guess the model” quiz at firmulate.com/quiz.html — a chance to see whether you can tell a 95-point operator from a 73-point one by its choices alone. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The takeaway for anyone evaluating AI tools — for business or for the workshop — is that “reads your files before answering” is not a soft quality. It’s a measurable, purchase-deciding property, and it’s worth real money: in this experiment, €55,000 up front and €4,583 a month, decided entirely by whether the agent did its homework two references deep.

So before you trust an AI with your CRM, your support queue, or your cut list, ask the woodworker’s question: does it measure twice? The league table says some do — and the full results and plain-language findings are published at firmulate.com/benchmarks.html. The board doesn’t care how good your saw is if you never checked the mark.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Why Cobalt Drill Bits Matter on Harder Metals

Kobalt drill bits matter on harder metals because their heat resistance and durability enable faster, more efficient drilling—discover how proper techniques can maximize their performance.

How to Choose a Circular Saw Blade for Clean Plywood Cuts

The key to clean plywood cuts lies in selecting the right circular saw blade, and understanding the factors that influence cut quality will help you achieve professional results.

Why Oscillating Tool Blades Wear So Differently

AIThis post was created with the assistance of artificial intelligence (AI).Your oscillating…

What Makes a Good Stop Block System for Repetitive Cuts

Inevitably, choosing the right stop block system ensures precision and safety, but discovering the key factors can significantly improve your setup.